故意找錯練習Spot the Error
本頁把 askLLM/docs/LIMITATIONS.zh-TW.md 裡記錄的錯誤與示警,改寫成一組「找碴」練習:你會先看到一段情境與一段 AI 回覆,回覆裡藏著一個錯誤,練習先自己找出問題所在,再打開答案核對。這是練習 verify/workflow.qmd 裡 V1(路徑核對)、 V3(範圍核對)與 V4a(讀碼核對)的具體方式——目的是讓你親身發現幻覺,而不是被動接受文件裡的結論。
This page turns the errors and warnings recorded in askLLM/docs/LIMITATIONS.zh-TW.md into a “spot the error” practice set: you will see a scenario and an AI reply with an error hidden in it, try to find the problem yourself, then open the answer to check. This is a hands-on way to practice V1 (path check), V3 (scope check), and V4a (code-read check) from verify/workflow.qmd. The point is to discover the hallucination yourself rather than take the documentation’s word for it.
題目來源說明:第 1–6 題逐字取材自 askLLM/docs/LIMITATIONS.zh-TW.md 表格中記錄的實測錯誤路徑,出處已在每題標明。第 7 題到第 10 題都是示意範例,題目本身會清楚標示。第 7 題改寫自 LIMITATIONS〈綜合使用建議〉第 3 點,該處的「8 個樣本、Saguaro/Palo Verde/Ironwood」是為說明錯誤型態而寫的虛構例子, LIMITATIONS 原文已明文標示它不是實測記錄。第 8–10 題是本頁自行設計,用來練習 V4a 的程式碼紅旗核對,依 askLLM/R/llm-adapter.R 明文禁止的指令清單改寫成教學用例。這四題都不是任何 AI 的真實輸出。第 11 題取自本站自己的實測記錄(data/ 底下的 .omv 與 data/replies/ 的回覆全文),第 12–14 題亦然。目前共 14 題,本頁會在課堂驗證後補進新的實測案例持續擴充。
第 1–6 題與第 11–14 題的 AI 回覆逐字引用自實測記錄,未經任何文字修飾,因此其標點與句式不遵循本站的散文規範(Pinker Rules 7–10);第 7–10 題的示意回覆則依這套規範改寫。
Where the questions come from: Questions 1–6 are taken verbatim from the real error paths recorded in the tables of askLLM/docs/LIMITATIONS.zh-TW.md; the source is noted on each question. Questions 7 through 10 are illustrative examples, and each is clearly labeled. Question 7 is adapted from point 3 of the “overall recommendations” in LIMITATIONS, where “8 samples across Saguaro / Palo Verde / Ironwood” is a made-up example written to show the failure mode. LIMITATIONS itself now says plainly that it is not a recorded test result. Questions 8–10 are designed for this page to practice the V4a code red-flag check, built from the instruction list that askLLM/R/llm-adapter.R explicitly forbids. None of these four is real output from any AI. Questions 11 through 14 come from this site’s own test records, the .omv files under data/ and the full replies in data/replies/. There are 14 questions here today, and the set will grow as classroom use turns up new cases.
The AI replies in questions 1–6 and questions 11 through 14 are quoted verbatim from the real-test record, with no stylistic editing, so their punctuation and sentence patterns do not follow this site’s prose style guide (Pinker Rules 7–10); the illustrative replies in questions 7–10 have been rewritten to follow it.
第 2 題:相關性長在哪裡Question 2: where correlation lives
來源:askLLM/docs/LIMITATIONS.zh-TW.md 第 1 節實測記錄。
情境:使用者問「我想看兩個連續變項的相關係數,該用哪個分析?」
AI 回覆:「請用 探索 > 相關性。」
問題:這條路徑錯在哪裡?
Source: recorded observation in section 1 of askLLM/docs/LIMITATIONS.zh-TW.md.
Scenario: A user asks “I want the correlation between two continuous variables — which analysis?”
AI reply: “Use Exploration > Correlation.”
Question: What is wrong with this path?
答案:相關矩陣實際位於 Regression 選單之下,不在 Exploration。跟第 1 題一樣,AI 正確辨識出「該做相關分析」,但把它掛在錯誤的頂層選單下。對應查核步驟:V1 路徑核對。
Answer: the correlation matrix actually lives under the Regression menu, not Exploration. As in question 1, the AI correctly identified “this calls for a correlation analysis” but hung it under the wrong top-level menu. Verification step: V1, path check.
第 3 題:主成分分析躲在哪裡Question 3: where PCA hides
來源:askLLM/docs/LIMITATIONS.zh-TW.md 第 1 節實測記錄。
情境:使用者問「我想對這幾個連續變項做主成分分析。」
AI 回覆:「可以到 探索 > 主成分分析。」
問題:這條路徑錯在哪裡?
Source: recorded observation in section 1 of askLLM/docs/LIMITATIONS.zh-TW.md.
Scenario: A user asks “I want to run a principal component analysis on these continuous variables.”
AI reply: “Go to Exploration > Principal Component Analysis.”
Question: What is wrong with this path?
答案:PCA 實際位於 Factor 選單之下,不在 Exploration。同一個錯誤模式再出現一次:分析名稱對,選單掛錯地方。對應查核步驟:V1 路徑核對。
Answer: PCA actually lives under the Factor menu, not Exploration. The same error pattern shows up again: the analysis name is right, the menu it’s hung under is wrong. Verification step: V1, path check.
第 5 題:編出來的巢狀層級Question 5: an invented nesting
來源:askLLM/docs/LIMITATIONS.zh-TW.md 第 1 節實測記錄。
情境:使用者用英文問 AI 如何做線性迴歸。
AI 回覆:「Use Regression > General Linear Model > Linear Regression.」
問題:這條路徑錯在哪裡?
Source: recorded observation in section 1 of askLLM/docs/LIMITATIONS.zh-TW.md.
Scenario: A user asks the AI, in English, how to run a linear regression.
AI reply: “Use Regression > General Linear Model > Linear Regression.”
Question: What is wrong with this path?
答案:LIMITATIONS 記錄這是「巢狀層級是編的」——Linear Regression 確實在 Regression 選單之下,但中間多插了一層不存在的 General Linear Model 子選單。這種錯誤特別難抓,因為頭尾都對,只有中間一層是編的;逐條核對每一層(不是只看第一層與最後一層對不對)才抓得到。對應查核步驟:V1 路徑核對。
Answer: LIMITATIONS records this as “a fabricated nesting level.” Linear Regression really is under Regression, but an extra, nonexistent General Linear Model sub-menu has been inserted in between. This kind of error is especially hard to catch because the first and last segments are both correct. Only the middle level is invented. You only catch it by checking every level of the path, not just whether the first and last segments look right. Verification step: V1, path check.
第 6 題:連翻譯都出錯Question 6: even the translation is wrong
來源:askLLM/docs/LIMITATIONS.zh-TW.md 第 1 節實測記錄。
情境:使用者用中文問 AI 如何做線性迴歸。
AI 回覆:「請到 迴歸 → 線性線性(Linear Regression)。」
問題:除了路徑本身,這則回覆還有什麼地方看起來不對勁?
Source: recorded observation in section 1 of askLLM/docs/LIMITATIONS.zh-TW.md.
Scenario: A user asks the AI, in Chinese, how to run a linear regression.
AI reply: “Go to 迴歸 → 線性線性(Linear Regression)” (the Chinese label literally reads “linear-linear”).
Question: Besides the path itself, what else looks off about this reply?
答案:LIMITATIONS 特別點出「連中文標籤都出現錯字」——線性線性 是重複詞,不是任何 jamovi 介面上真實出現過的字串。這提醒我們:路徑核對不能只看「大致講得通」,連標籤的逐字拼寫都要對得上你自己 jamovi 裡看到的文字,因為自信的語氣與正確性無關。對應查核步驟:V1 路徑核對。
Answer: LIMITATIONS specifically flags that “even the Chinese label has a typo.” 線性線性 (“linear-linear”) is a duplicated word that never appears anywhere in jamovi’s actual interface. This is a reminder that the path check isn’t satisfied by something that “roughly makes sense.” The label’s exact wording has to match what you see in your own jamovi. Confident tone has nothing to do with correctness. Verification step: V1, path check.
第 7 題(示意範例):沒附摘要時的憑空生成Question 7 (illustrative): fabrication with no summary attached
本題為示意範例,並非任何 AI 的真實輸出。LIMITATIONS 用這個虛構例子說明「沒有資料時模型會編一份出來」這種錯誤型態。
來源:askLLM/docs/LIMITATIONS.zh-TW.md〈綜合使用建議〉第 3 點的示意範例。
情境:使用者在送出問題時沒有勾選「Attach data summary to prompt」,資料集其實是 iris(150 筆鳶尾花觀測,三個品種:setosa、versicolor、 virginica),問 AI「這份資料有幾個樣本、有哪些組別?」
AI 回覆(示意):「這份資料大約有 8 個樣本,分成 Saguaro、Palo Verde 與 Ironwood 三個物種。」
問題:這則回覆錯在哪裡?根本原因是什麼?
This question is an illustrative example, not real output from any AI. LIMITATIONS uses this made-up case to show the failure mode: given no data, the model makes some up.
Source: the illustrative example in the “overall recommendations” section of askLLM/docs/LIMITATIONS.zh-TW.md, point 3.
Scenario: A user submits a question without ticking “Attach data summary to prompt.” The dataset is actually iris (150 observations, three species: setosa, versicolor, virginica), and the user asks “how many samples are in this dataset, and what groups does it have?”
AI reply (illustrative): “This dataset has roughly 8 samples, split into three species: Saguaro, Palo Verde, and Ironwood.”
Question: What is wrong with this reply? What is the root cause?
答案:這是 LIMITATIONS 用來說明錯誤型態的示意範例。未勾選「Attach data summary」時,AI 完全看不到你的資料,卻仍然給出具體數字(8 個樣本)與具體物種名稱(Saguaro、Palo Verde、Ironwood,這三者都不是 iris 的品種,是憑空生成的仙人掌/植物名稱)。根本原因不是「AI 算錯了」,是它從頭到尾沒有拿到任何資料資訊,卻用跟有資料時同樣肯定的語氣回答。對應查核步驟:V3 範圍核對 ——回覆聲稱知道它其實看不到的東西,就要整段打折扣,並回頭確認送出問題時是否真的附上了資料摘要。
Answer: LIMITATIONS uses this made-up case to show the failure mode. With “Attach data summary” unticked, the AI has no visibility into your data at all, yet it still produces a specific number (8 samples) and specific group names. Saguaro, Palo Verde, and Ironwood are not iris species at all. They are fabricated plant names. The root cause is not “the AI miscalculated.” It never received any information about the data in the first place, yet answered with the same confident tone it would use if it had. Verification step: V3, scope check. Discount entirely any reply claiming to know something it cannot actually see, then go back and confirm whether the data summary was actually attached when you submitted the question.
第 8 題(示意範例):程式碼裡的安裝指令Question 8 (illustrative): an install call in the code
本題為示意範例,並非任何 AI 的真實輸出,用來練習 V4a 讀碼核對。
情境:使用者用 R code tutor 問「幫我對 data 裡的 score 欄位做常態性檢定。」
AI 回覆(示意):
install.packages("moments")
library(moments)
agostino.test(data$score)問題:這段程式碼違反了哪一條紅旗規則?
This question is an illustrative example, not real output from any AI, used to practice the V4a code-read check.
Scenario: A user asks the R code tutor “run a normality test on the score column in data.”
AI reply (illustrative):
install.packages("moments")
library(moments)
agostino.test(data$score)Question: Which red-flag rule does this code violate?
答案:第一行 install.packages("moments") 直接違反 askLLM/R/llm-adapter.R 第 172–177 行明文禁止的指令清單。System prompt 明確要求「never suggest install.packages()」。這行不能執行。你的 Rj 環境不保證有網路權限或寫入權限去安裝新套件。askLLM 應該只引用 <rj_environment> 清單裡已經裝好的套件。對應查核步驟:V4a 讀(見 checklist-rtutor.qmd 的紅旗字串表)。
Answer: askLLM/R/llm-adapter.R lines 172–177 explicitly forbid this instruction. The first line, install.packages("moments"), violates it directly. The system prompt states plainly: never suggest install.packages(). This line cannot run safely. Your Rj environment offers no guarantee of network access or write access to install a new package. askLLM must cite only packages already installed, from the <rj_environment> list. Verification step: V4a, read (see the red-flag table in checklist-rtutor.qmd).
第 9 題(示意範例):程式碼裡的檔案讀取Question 9 (illustrative): a file read in the code
本題為示意範例,並非任何 AI 的真實輸出,用來練習 V4a 讀碼核對。
情境:使用者用 R code tutor 問「幫我畫出兩個變項的散佈圖。」
AI 回覆(示意):
mydata <- read.csv("C:/Users/me/Desktop/survey.csv")
plot(mydata$x, mydata$y)問題:這段程式碼違反了哪幾條紅旗規則?(不只一條)
Scenario: A user asks the R code tutor “plot a scatterplot of two variables.”
AI reply (illustrative):
mydata <- read.csv("C:/Users/me/Desktop/survey.csv")
plot(mydata$x, mydata$y)Question: Which red-flag rules does this code violate? (More than one.)
答案:這一行同時踩到兩條紅旗。第一條是 read.csv(...):程式碼不得建議讀取檔案,資料一律要用既有的 data 物件。第二條是絕對路徑: C:/Users/me/Desktop/... 這種路徑因人而異,在你的機器上幾乎必然不存在。你的桌面上也許剛好有一份叫 survey.csv 的檔案,但這段程式碼依然完全沒有用到 jamovi 已經幫你準備好的 data 物件,等於繞過了 askLLM 的設計邊界。對應查核步驟:V4a 讀。
Answer: this single line trips two red flags. The first is read.csv(...): no file reads are allowed, and the data must always come from the existing data object. The second is the absolute path: C:/Users/me/Desktop/... is machine-specific and almost certainly does not exist on your computer. Your desktop might happen to hold a file called survey.csv. This code still never touches the data object jamovi already prepared for you, and it sidesteps askLLM’s design boundary. Verification step: V4a, read.
第 10 題(示意範例):引用一個沒裝的套件Question 10 (illustrative): citing an uninstalled package
本題為示意範例,並非任何 AI 的真實輸出,用來練習 V4a 讀碼核對。
情境:使用者的 Rj 環境裡只裝了 base R(沒有額外安裝任何統計套件), <rj_environment> 清單因此只列出 base R 套件。使用者問「幫我做一個穩健迴歸(robust regression)。」
AI 回覆(示意):
library(MASS)
fit <- rlm(y ~ x, data = data)
summary(fit)問題:這段程式碼有問題嗎?如果 MASS 沒有出現在 <rj_environment> 清單裡,你會怎麼處理?
Scenario: The user’s Rj environment has only base R installed (no extra statistical packages), so the <rj_environment> list contains only base R packages. The user asks “run a robust regression for me.”
AI reply (illustrative):
library(MASS)
fit <- rlm(y ~ x, data = data)
summary(fit)Question: Is there a problem with this code? If MASS does not appear in the <rj_environment> list, what should you do?
答案:這裡要看清楚:MASS 隨 base R 一起發行,在多數 R 安裝裡本來就存在,所以「引用 MASS」本身不一定是錯的。關鍵是核對,不是背名單。真正的檢查方法是:打開這次呼叫附給 AI 的 <rj_environment> 區塊,逐字確認 MASS 是否真的列在裡面。如果列在裡面,這段程式碼沒有違反紅旗規則;如果沒列在裡面(代表你的 Rj 掃描結果不含它),這段程式碼一執行就會得到 there is no package called 'MASS'。這種情況下應該回頭修改問題(例如「只用 base R 做穩健估計」),而不是自己手動安裝套件來配合 AI 的回覆。對應查核步驟:V4a 讀 與 V4b 跑(V4b 的錯誤訊息正是拿來確認 V4a 判斷是否正確的依據)。
Answer: look carefully here. MASS ships with base R and exists in most R installations by default, so citing MASS is not automatically wrong. The point is to check, not to memorize a blocklist. The real check is: open the <rj_environment> block sent to the AI for this call, and confirm verbatim whether MASS is actually listed. If it is, this code does not violate a red flag. If it is not listed, meaning your Rj scan didn’t find it, running this code will immediately produce there is no package called 'MASS'. In that case, go back and rephrase the question, for example “use only base R for a robust estimate,” rather than manually installing a package just to make the AI’s reply work. Verification steps: V4a, read and V4b, run (the V4b error message is exactly what confirms whether the V4a judgment was correct).
第 11 題:同一個模型,兩種問法Question 11: same model, two ways of asking
本題取自
data/mixed-anova-setup_zhang2014_v1_fail_zhTW.omv的實測記錄,回覆逐字引用,未經修飾。
情境:使用者請 R code tutor 幫忙把寬格式資料轉成長格式,並在提示詞裡要求「只給程式骨架與 TODO 註解」。模型是 gemini-flash-latest。
AI 回覆(節錄):
data_long <- data %>%
pivot_longer(
cols = c(# TODO: 填入要轉換的兩個時間點欄位名稱),
names_to = "# TODO: 指定時間點的新欄位名稱 (例如 'Time')",
values_to = "# TODO: 指定測量分數的新欄位名稱 (例如 'Interest')"
)問題:這段程式碼貼進 Rj 會發生什麼事?錯在模型,還是錯在問法?
This question comes from the recorded test in
data/mixed-anova-setup_zhang2014_v1_fail_zhTW.omv. The reply is quoted verbatim, with no editing.
Scenario: The user asked the R code tutor to reshape wide data into long format, and the prompt asked for “only a skeleton with TODO comments.” The model was gemini-flash-latest.
AI reply (excerpt):
data_long <- data %>%
pivot_longer(
cols = c(# TODO: fill in the two time-point column names),
names_to = "# TODO: name for the new time column",
values_to = "# TODO: name for the new score column"
)Question: What happens when this is pasted into Rj? Is the fault the model’s, or the question’s?
答案:貼進 Rj 會立刻得到 <text>:23:1: unexpected symbol,程式碼一行都跑不了。原因是 R 的 # 會讓該行剩下的所有內容成為註解——包括 c( 需要的那個右括號。運算式因此永遠收不了尾,解析器在後面第一個非空白符號處報錯。英文版的同一次實測更嚴重,還出現 table(data_long$# TODO: ...)。
錯的是問法。 提示詞說「只給程式骨架與 TODO 註解」,等於邀請模型把註解填在該放值的位置。同一個模型、同一份資料、同一個 persona、同一種提示詞語言,只把那一句改成「待填處寫成開頭的變數指派、值放字串佔位符」之後,產出就變成:
target_cols <- c("TODO_填入T1欄位名稱", "TODO_填入T2欄位名稱")
time_col_name <- "Time"
data_long <- data %>%
pivot_longer(cols = all_of(target_cols), names_to = time_col_name, ...)待填處還在,但它們是字串,程式碼在你填值之前就能通過語法檢查。實測結果從 fail 變成 pass(見提示詞庫該條目的 tested_with)。
這一題真正要示範的不是那個 bug,而是查核方法本身:兩次執行只差一個因子,所以差異可以歸因。這正是 workflow.qmd 的 cross-check 與第六課「換模型或換 persona 重問一次」在教的事。如果那兩次同時換了模型又換了問法,你就無法說出是哪一個造成差異。
對應查核步驟:V4a 讀 與 V4b 跑。這也是為什麼 V4b 不能省略——這段程式碼「讀起來」很合理,錯誤只有執行才會現形。
Answer: pasting this into Rj immediately produces <text>:23:1: unexpected symbol, and not a line of it runs. R’s # turns everything remaining on that line into a comment, including the closing bracket that c( needs. The expression never terminates, and the parser reports the error at the first non-blank symbol after it. The English run of the same test was worse still, producing table(data_long$# TODO: ...).
The fault is the question’s. The prompt asked for “only a skeleton with TODO comments,” which invites the model to put comments where values belong. Same model, same data, same persona, same prompt language: change that one sentence to “put every blank as a variable assignment at the top, with a string placeholder as its value,” and the output becomes:
target_cols <- c("column name here", "column name here")
time_col_name <- "Time"
data_long <- data %>%
pivot_longer(cols = all_of(target_cols), names_to = time_col_name, ...)The blanks are still there, but they are strings, so the code parses before you fill anything in. The recorded result went from fail to pass (see the entry’s tested_with in the prompt library).
What this question really demonstrates is not the bug but the verification method. The two runs differed by exactly one factor, so the difference can be attributed. That is what the cross-check in workflow.qmd and Lesson 6’s “ask again with a different model or persona” are teaching. Had those two runs changed the model and the wording, you could not say which one caused the difference.
Verification steps: V4a, read and V4b, run. This is also why V4b cannot be skipped: the code reads perfectly sensibly, and the error appears only when it runs.
第 12 題:每一句都是真的Question 12: every sentence is true
本題取自
data/paired-vs-independent_stroop_zhTW.omv的實測記錄,回覆逐字引用,未經修飾。回覆全文見data/replies/paired-vs-independent_stroop_zhTW.md。
情境:使用者把 Stroop 資料的三個變項勾進 askLLM,提問裡寫明「同一位受試者在兩個水準下都有資料」,然後問:「這個資料是配對的。你從我提供的摘要統計裡看得出配對關係嗎?請誠實回答。」模型是 gemini-flash-latest, persona 是預設的 consultant。
AI 回覆(第 2 題部分):
是否看得出配對關係:
可以看出高度配對的特徵。從摘要統計中可清楚看到:
participant_id共有 270 個水準(受試者),且每位受試者皆正好出現 2 次(計 540 筆資料)。condition的兩個水準(congruent 與 incongruent)剛好各為 270 筆。這與「每位受試者在兩種條件各測量一次」的配對設計完全吻合。(誠實補充:摘要中未提供交叉列聯表,因此技術上無法 100% 排除極少數受試者重複錄入同一條件的例外,但結構上完全呈現配對特徵)。
問題:拿資料去查,這段話裡的每一件事實都正確。那它為什麼仍然該被標為越界?
This question comes from the recorded test in
data/paired-vs-independent_stroop_zhTW.omv. The reply is quoted verbatim, with no editing. The full reply is indata/replies/paired-vs-independent_stroop_zhTW.md.
Scenario: The user ticked three Stroop variables into askLLM. The prompt stated that the same participant has data under both levels, then asked: “This data is paired. Can you tell that from the summary statistics I sent? Answer honestly.” The model was gemini-flash-latest and the persona was the default consultant.
AI reply (the part answering question 2, translated from the Chinese run):
Can I see the pairing?
Yes, the summary shows strong signs of pairing:
participant_idhas 270 levels, one per participant, and each participant appears exactly twice, for 540 rows in total.- The two levels of
condition, congruent and incongruent, hold exactly 270 rows each.This matches a paired design in which every participant is measured once under each condition. (In fairness: the summary gives no cross-tabulation, so strictly speaking I cannot rule out a handful of participants recorded twice under the same condition. Structurally, though, the pairing is fully evident.)
Question: Check it against the data and every factual statement here is correct. So why should this still be marked as overreaching?
答案:因為它拿不到支撐那句話的證據。
摘要統計給模型的是「participant_id 有 270 個水準」與「總列數 540」兩個數字。由 540 ÷ 270 = 2 推不出「每位受試者皆正好出現 2 次」——一位受試者 3 列、另一位 1 列,算出來的平均一模一樣。要確認每人各 2 列,需要的是每位受試者的列數分配,而摘要裡沒有這個東西。
那句話碰巧是真的。實算確認資料確實是 270 人、每人正好 2 列、每人每條件正好 1 筆(270/270)。但結論正確不等於推理成立。模型下次在另一份資料上做同樣的推理,就會錯,而你沒有任何辦法從回覆本身分辨這兩種情況。
它括號裡的補充其實說對了:「摘要中未提供交叉列聯表,因此技術上無法 100% 排除 ⋯⋯」。問題是那句話被放在括號裡,標題已經寫成「可以」。初學者讀到的是粗體的「可以」,不是括號裡的但書。
對照組是現成的。 同一條提示詞的英文版,同模型、同資料、同 persona,只差 promptLang,第 2 題答的是:
Partially, but not definitively. [⋯] the summary does not include a cross-tabulation of
participant_idbycondition. Therefore, it cannot strictly verify that each participant has one congruent and one incongruent observation rather than, for example, two observations in the same condition.
同樣的認識論限制,一個放進標題、一個塞進括號。兩次記錄分別是 pass 與 partial(見提示詞庫該條目的 tested_with)。
配對與否是研究設計決定的,不是資料長相決定的。這件事只有你知道,模型必須由你告知。
對應查核步驟:V3 範圍核對——回覆有沒有宣稱它「看不到」的事。這一題示範 V3 為什麼不能只查事實對錯:事實全對,仍然不通過。順帶示範 V5 交叉,兩次執行只差一個因子,差異因此可以歸因。
Answer: because it had no evidence for what it said.
What the summary handed the model was two numbers: participant_id has 270 levels, and there are 540 rows. From 540 ÷ 270 = 2 you cannot derive “each participant appears exactly twice.” One participant with 3 rows and another with 1 gives the same average. Establishing two rows each needs the distribution of rows per participant, and the summary does not carry it.
The claim happens to be true. The data really does hold 270 participants with exactly 2 rows each, and exactly one row per condition, 270 out of 270. But a correct conclusion is not a sound inference. Run the same reasoning on the next dataset and it will be wrong, and nothing in the reply itself lets you tell the two cases apart.
Its parenthetical got this right: the summary gives no cross-tabulation, so strictly speaking it cannot rule out the alternative. The trouble is that the caveat sits in a parenthesis while the heading says Yes. A beginner reads the bold word, not the bracket.
The control run already exists. The English version of the same entry, same model, same data, same persona, differing only in promptLang, answered:
Partially, but not definitively. […] the summary does not include a cross-tabulation of
participant_idbycondition. Therefore, it cannot strictly verify that each participant has one congruent and one incongruent observation rather than, for example, two observations in the same condition.
The same epistemic limit went into the headline in one run and into a parenthesis in the other. The two records are pass and partial (see the entry’s tested_with in the prompt library).
Whether data is paired is settled by the study design, not by what the data looks like. Only you know it, and you have to tell the model.
Verification step: V3, scope check, does the reply claim something it cannot see. This question shows why V3 cannot stop at checking facts: every fact is right and it still fails. It also shows V5, cross-check, where two runs differ by one factor so the difference can be attributed.
第 13 題:模型給了一個「照做就對得上」的錯誤指示Question 13: an instruction that looks like it would match, but doesn’t
本題取自
data/describe-code-crosscheck_stroop_en.omv與data/describe-code-crosscheck_stroop_zhTW.omv的實測記錄,回覆逐字引用,未經修飾。回覆全文見data/replies/。
情境:使用者請 R code tutor 寫一段描述統計程式碼,要拿去跟 jamovi Descriptives 逐格比對,並請它說明哪些統計量可能在 R 與 jamovi 之間算不出同一個數字。資料是 Stroop(270 位受試者 × congruent/incongruent)。
AI 回覆(英文版,節錄):
Quartiles (25th and 75th percentiles): Base R’s
quantile()defaults totype = 7, whereas jamovi usestype = 6(standard in SPSS/Minitab). To match jamovi, you must specifytype = 6in R’squantile()function.
中文版回覆說的是同一件事:「R 的 quantile() 預設演算法是 type = 7;而 jamovi (以及 SPSS) 預設計算四分位數採用的是 type = 6(以 \(p(N+1)\) 定位)。」同一回覆對偏態峰度的說明是對的:「jamovi reports sample-adjusted values (Type 2, equivalent to SPSS/SAS)」。
問題:這段關於四分位數的指示錯在哪裡?跟前面幾題的錯誤比起來,這種錯法為什麼特別危險?
This question comes from the recorded tests in
data/describe-code-crosscheck_stroop_en.omvanddata/describe-code-crosscheck_stroop_zhTW.omv. The reply is quoted verbatim, with no editing. Full replies are indata/replies/.
Scenario: A user asked the R code tutor to write descriptive-statistics code to check cell by cell against jamovi’s Descriptives, and to note which statistics might not come out the same between R and jamovi. The dataset was Stroop data, 270 participants under congruent and incongruent conditions.
AI reply (English run, excerpt):
Quartiles (25th and 75th percentiles): Base R’s
quantile()defaults totype = 7, whereas jamovi usestype = 6(standard in SPSS/Minitab). To match jamovi, you must specifytype = 6in R’squantile()function.
The Chinese run said the same thing: 「R 的 quantile() 預設演算法是 type = 7;而 jamovi (以及 SPSS) 預設計算四分位數採用的是 type = 6 (以 \(p(N+1)\) 定位)。」 The same reply’s note on skewness and kurtosis was correct: “jamovi reports sample-adjusted values (Type 2, equivalent to SPSS/SAS).”
Question: What is wrong with this quartile instruction? Compared with the errors in earlier questions, why is this particular kind of mistake more dangerous?
答案:錯的是那句「To match jamovi, you must specify type = 6」。jamovi 的 Descriptives 實際上用的是 R 的預設值 type = 7——jamovi 28.2 內建的 jmv 2.8.0 呼叫 quantile() 時根本沒有指定 type 參數。照模型的指示加上 type = 6,程式碼會順利執行、數字看起來也合理,但四分位數對不上 jamovi。實測對照:
| jamovi Descriptives | Rj 照建議用 type = 6 |
|
|---|---|---|
| congruent Q1/Q3 | 667.93/818.76 | 667.48/819.07 |
| incongruent Q1/Q3 | 796.78/984.23 | 796.22/984.68 |
n、mean、sd、median、min、max 六項完全一致,只有四分位數全部不符。
這一題要示範的重點:
- 模型把一個錯的事實包裝成「照做就能對上」的具體指示,比一般的模糊說法更危險——它給了你一個看似可驗證、可執行的動作,讓你更不會去懷疑。
- 這個錯誤在中英兩版、以及提示詞補完整後重跑的版本都出現,是穩定的錯,不是偶發的幻覺。
- 抓到它的是 cross-check 這一步:只看 Rj 輸出不會發現任何問題,程式碼零錯誤、數字看起來合理,問題只有拿去跟 jamovi 逐格比對才會現形。
- 小數點後零點幾的差距很容易被當成捨入誤差放過,需要提醒自己:捨入誤差不會系統性地只出現在四分位數這一類統計量上。
對應查核步驟:V5 交叉(與 jamovi 輸出逐格比對)。
Answer: the broken part is “To match jamovi, you must specify type = 6.” jamovi’s Descriptives actually uses R’s own default, type = 7. jamovi 28.2’s bundled jmv 2.8.0 calls quantile() without ever setting type. Follow the model’s instruction and add type = 6, and the code runs fine and the numbers look reasonable, but the quartiles do not match jamovi. The recorded comparison:
| jamovi Descriptives | Rj with the suggested type = 6 |
|
|---|---|---|
| congruent Q1/Q3 | 667.93/818.76 | 667.48/819.07 |
| incongruent Q1/Q3 | 796.78/984.23 | 796.22/984.68 |
n, mean, sd, median, min, and max all match exactly. Only the quartiles are off, every single one of them.
What this question demonstrates:
- Packaging a wrong fact as a specific, “do this and it will match” instruction is more dangerous than a vague claim. It hands you a verifiable-looking action. That makes you less likely to doubt it.
- The error showed up in both the Chinese and English runs, and again after the prompt was filled out and rerun. It is a stable error, not a one-off hallucination.
- Cross-checking is what catches it. Reading the Rj output alone finds nothing wrong: the code runs with zero errors and the numbers look plausible. The mismatch only shows up once you compare cell by cell against jamovi.
- A difference of a few tenths in the decimals is easy to wave off as rounding error. Rounding error would not hit all four quartiles and leave the other six statistics exact.
Verification step: V5, cross-check (compare cell by cell against jamovi’s own output).
第 14 題:選項是真的,位置是編的Question 14: the option is real, its location is not
本題取自
data/three-group-comparison_monin2008_zhTW.omv的實測記錄,回覆逐字引用,未經修飾。對照組見同一條目的data/three-group-comparison_monin2008_en.omv。
情境:使用者問 Module Guider 比較三組平均數該用哪個分析、有哪些前提要檢查、各自在 jamovi 哪裡勾選。資料是 RP:P #43 Monin 複製研究,Condition 分三組(20/27/28 人),Intelligent 為 1–7 評分。
AI 回覆(中文版,節錄):
在 jamovi 中的勾選位置: 進入
Analyses > ANOVA > One-Way ANOVA分析設定面板後: * 在 Assumption Checks 區塊中: * 勾選 Normality (Shapiro-Wilk)(檢驗常態性) * 勾選 Q-Q plot(視覺化輔助檢視常態性) * 在 Variances 區塊中: * 可勾選 Homogeneity (Levene’s) 檢定變異數同質性;若同質性不成立,可直接在同區塊中選取 Don’t assume equal (Welch’s) 進行校正。
問題:對照你自己 jamovi 裡的 One-Way ANOVA 面板,這段勾選位置錯在哪裡?
This question comes from the recorded test in
data/three-group-comparison_monin2008_zhTW.omv. The reply is quoted verbatim, with no editing. The control run is the English version of the same entry,data/three-group-comparison_monin2008_en.omv.
Scenario: A user asked the Module Guider which analysis compares three group means, what assumptions to check, and where each check lives in jamovi. The dataset was the RP:P #43 Monin replication, with Condition split into three groups (20, 27, 28 participants) and Intelligent rated 1–7.
AI reply (Chinese run, excerpt, translated here; the original Chinese is quoted in the [zh] version of this question above):
Where to check these in jamovi: Open the
Analyses > ANOVA > One-Way ANOVAsettings panel, then: * Under Assumption Checks: * Check Normality (Shapiro-Wilk) * Check Q-Q plot * Under Variances: * You may check Homogeneity (Levene’s) to test for equal variances; if it fails, select Don’t assume equal (Welch’s) in the same section to correct for it.
Question: Checked against your own jamovi’s One-Way ANOVA panel, what is wrong with these locations?
答案:jamovi 的 One-Way ANOVA 面板裡,Variances 區塊只有兩個選項: Don’t assume equal (Welch’s) 與 Assume equal (Fisher’s)。同質性檢定並不在這裡,它跟 Normality test、Q-Q plot 一起放在 Assumption Checks 區塊。模型把 Homogeneity 檢定挪到了 Variances 底下,編出一個它自己聽起來很合理、但面板上不存在的排列。
選項名稱本身也不是逐字的:實際介面上寫的是 Homogeneity test 與 Normality test,不是模型寫的 Homogeneity (Levene’s) 與 Normality (Shapiro-Wilk)——後面那個帶括號的方法名稱是模型自己加上去的說明文字,不是 jamovi 介面上出現的字串。
Welch’s 校正的位置模型倒是說對了:它確實在 Variances 區塊,跟 Assume equal (Fisher’s) 並列。
這一題要示範的重點:
- 選項本身存在、Welch’s 的位置也對,唯獨 Homogeneity test 被搬到了錯的區塊——這種「大致正確、細節編造」比第 4 題那種整條路徑純虛構更難察覺,因為你光看「有沒有這個選項」會覺得沒問題。
- path 核對要點到最後一層:不能只確認分析名稱(One-Way ANOVA)對、大選單(Analyses > ANOVA)對,還要逐一核對子區塊底下每個選項究竟被歸在哪一區。
- 與第 12 題呼應:同一條提示詞的英文版把三個選項都正確地寫在 Assumption Checks 底下(“Expand the Assumption Checks section and check Homogeneity test (Levene’s test)”)。同一個模型、同一份資料、同一個 persona,只差
promptLang,可靠度不一樣。
對應查核步驟:V1 路徑核對(核對到子區塊層級,不是只核對選單與分析名稱)。
Answer: jamovi’s One-Way ANOVA panel has only two options under Variances: Don’t assume equal (Welch’s) and Assume equal (Fisher’s). The homogeneity test is not there. It sits under Assumption Checks, alongside the normality test and the Q-Q plot. The model moved the homogeneity test into Variances. That layout sounds plausible, but it does not exist on the panel.
The option labels themselves are not verbatim either. The actual interface reads Homogeneity test and Normality test, not the model’s Homogeneity (Levene’s) and Normality (Shapiro-Wilk). The method names in parentheses are the model’s own added gloss. jamovi’s interface does not use these strings anywhere.
The model did get one thing right: Welch’s correction really is under Variances, next to Assume equal (Fisher’s).
What this question demonstrates:
- The options themselves exist, and Welch’s location is right. Only the homogeneity test is filed under the wrong section. Just checking “does this option exist” looks fine, so this “mostly right, with one fabricated detail” pattern is harder to catch than question 4’s fully invented path.
- A path check has to go all the way down. It is not enough to confirm the analysis name (One-Way ANOVA) and the top menu (Analyses > ANOVA). Every option has to be checked against the sub-section where it actually sits.
- This echoes question 12. The English run of the same prompt correctly put all three options under Assumption Checks: “Expand the Assumption Checks section and check Homogeneity test (Levene’s test).” Same model, same data, same persona, differing only in
promptLang, and the reliability differs.
Verification step: V1, path check (checked down to the sub-section level, not just the menu and analysis name).