提示詞庫Prompt Library

本頁由 tools/build-prompts.R 依 prompts/entries/*.yaml 自動產生,請勿手動修改本頁內容—— 要修改條目請改對應的 yaml 檔後重新執行 build 腳本。

This page is generated by tools/build-prompts.R from prompts/entries/*.yaml. Do not edit this page by hand – edit the corresponding YAML entry and rerun the build script instead.

用 R 檢驗兩個類別變項的關聯,並自己讀結果Testing association between two categorical variables in R, and reading the result yourself

analysis: rtutor  |  design: none  |  stat_goal: associate  |  persona: tutor

情境Scenario

有兩個類別變項,想請 R 跑卡方檢定,程式碼要能讓自己檢查前提是否成立,也不要模型替你解讀結果。

You have two categorical variables and want R code for a chi-square test. The code should let you check the assumptions yourself, and you do not want the model to interpret the result for you.

送出前必做Prerequisites

  • 先確認兩個變項在 jamovi 裡都是 Nominal 型別,且都沒有遺漏值
  • Rj 不會顯示 R 的 warning(),程式碼印出的期望次數表要自己看,別等模型幫你抓
  • Confirm both variables are set to Nominal in jamovi, and that neither has missing values.
  • Rj does not display R’s warning() messages. The code prints an expected-count table; look at it yourself instead of waiting for the model to flag it.

提示詞Prompt

data 有 {var_a}(3 水準:Man、Woman、Non-binary)與 {var_b}(5 水準),{n} 列,兩欄都沒有遺漏。

請給我一段 R 程式碼檢驗這兩個類別變項是否有關聯,我之後要拿去和 jamovi 的結果比對。

規則:
1. 待填處寫成開頭的變數指派、值放字串佔位符,不要用註解當佔位符。
2. 只用 data,不要 read.csv、不要 install.packages。
3. 程式碼要讓我能自己檢查這個檢定的前提有沒有被滿足。
4. 不要替我解釋結果代表什麼,我要自己讀。
The data has {var_a} (3 levels: Man, Woman, Non-binary) and {var_b}
(5 levels), {n} rows, with no missing values in either column.

Please give me R code that tests whether these two categorical variables
are associated. I will compare the result against jamovi afterwards.

Rules:
1. Put every blank as a variable assignment at the top, with a string
   placeholder as its value. Do not use comments as placeholders.
2. Use data only. No read.csv, no install.packages.
3. The code should let me check for myself whether the test's assumptions
   hold.
4. Do not interpret the result for me. I want to read it myself.

期望回覆要素Expected elements

  • 只用 data,不出現 read.csv / install.packages(R tutor 的硬規則)
  • 待填處是開頭的變數指派、值為字串佔位符,填值前即可通過語法檢查
  • 程式碼會印出期望次數表(例如 chisq.test(...)$expected)
  • 沒有替使用者宣稱兩者有關聯或暗示有傾向;提示詞已要求不解釋
  • Uses only data, with no read.csv / install.packages (a hard rule for the R tutor)
  • Blanks are variable assignments at the top with string placeholders as their values; the code passes a syntax check before they are filled in
  • The code prints the expected-count table (for example chisq.test(...)$expected)
  • Does not claim the two variables are associated or hint at a trend for the user; the prompt already asked it not to interpret

查核點Check

  • code-read: 只用 base R(或 rj_environment 裡有的套件)、無 setwd/路徑/system
  • code-run: 貼進 Rj 執行,記錄卡方統計量、df、p 值,並看程式碼印出的期望次數表,自行檢查是否有格 < 5
  • cross-check: 回 jamovi 跑 Frequencies ▸ Independent Samples χ² test of association,比對 χ²、df、p 是否一致
  • code-read: Uses only base R (or packages available in rj_environment); no setwd, file paths, or system calls
  • code-run: Paste it into Rj and run it, recording the chi-square statistic, df, and p value, then look at the printed expected-count table yourself to check whether any cell is below 5
  • cross-check: Go back to jamovi and run Frequencies ▸ Independent Samples chi-square test of association, and compare the chi-square, df, and p against the R output

判準Stop criteria

  • solved: code-run 的統計量與期望次數表都核對過,cross-check 與 jamovi 結果一致,且回覆全程沒有替你解讀
  • reopen: 回覆把不顯著的 p 值講成「接近顯著」或「有傾向」、或完全沒提期望次數 → 逐字記下原文,換 consultant persona 重問一次比對
  • solved: The code-run statistics and the expected-count table have both been checked, the cross-check matches jamovi, and the reply never interpreted the result for you.
  • reopen: The reply describes a non-significant p value as “approaching significance” or “a trend”, or never mentions the expected counts -> record the text verbatim and ask again with the consultant persona to compare.

實測記錄Tested with

  • 2026-09-18 | gemini | gemini-flash-latest | pass | 補做 cross-check 後重評,zh 版,persona=tutor,promptLang=zh,證據 data/contingency-association_ballou2024_zhTW.omv、回覆副本 data/replies/contingency-association_ballou2024_zhTW.md。expected 4/4:只用 data;開頭兩個字串佔位符(var1/var2);印出 chisq.test 結果的 $expected;未替使用者解讀,也未把 p = .093 說成有傾向。check 三項齊備:code-read 只用 base R;code-run 貼進 Rj 的程式碼與回覆同源(只把佔位符換成 gender/eduLevel),輸出 X-squared = 13.594、df = 8、p-value = 0.09299,期望次數表含 4.875346,零錯誤;Rj 不顯示 warning,前提改由期望次數表自行判斷。cross-check:jamovi Frequencies ▸ Independent Samples χ² test of association 得 χ² = 13.59、df = 8、p = .093,與 Rj 一致。原判 partial(2026-09-17,缺 cross-check),補做後改判 pass。
  • 2026-09-18 | gemini | gemini-flash-latest | pass | 補做 cross-check 後重評,en 版,persona=tutor,promptLang 未設(en 預設),證據 data/contingency-association_ballou2024_en.omv、回覆副本 data/replies/contingency-association_ballou2024_en.md。expected 4/4:只用 data;開頭兩個字串佔位符(col_gender/col_education);印出 chisq.test 結果的 $expected;未替使用者解讀。check 三項齊備:code-read 只用 base R;code-run 貼進 Rj 的程式碼與回覆同源(只把佔位符換成 gender/eduLevel),輸出 X-squared = 13.594、df = 8、p-value = 0.09299,期望次數表含 4.875346,零錯誤。cross-check:jamovi Frequencies ▸ Independent Samples χ² test of association 得 χ² = 13.59、df = 8、p = .093,與 Rj 一致。原判 partial(2026-09-17,缺 cross-check),補做後改判 pass。

Full test notes are recorded in Chinese. Use the language toggle to read them.

  • 2026-09-18 | gemini | gemini-flash-latest | pass | Evidence: data/contingency-association_ballou2024_zhTW.omv, data/replies/contingency-association_ballou2024_zhTW.md
  • 2026-09-18 | gemini | gemini-flash-latest | pass | Evidence: data/contingency-association_ballou2024_en.omv, data/replies/contingency-association_ballou2024_en.md

兩個連續變項的關聯該用哪種相關Which correlation for two continuous variables

analysis: guider  |  design: none  |  stat_goal: associate  |  persona: consultant

情境Scenario

有兩個連續變項想看關聯,不確定該用 Pearson 還是 Spearman,也想知道哪些判斷需要看散佈圖。

You have two continuous variables and want to examine their association, but are unsure whether to use Pearson or Spearman, and which judgements need a scatterplot.

送出前必做Prerequisites

  • 兩個變項都勾進 Variables to describe,並記下各自的遺漏數
  • 先看 min 與 max,若兩者差距懸殊,之後要留意是否有極端值主導相關
  • Tick both variables into Variables to describe, and note how many values are missing in each.
  • Look at min and max first; a wide gap is a hint that extreme values may drive the correlation.

提示詞Prompt

{var_a} 與 {var_b} 都是連續變項,各有 {n_a} 與 {n_b} 筆有效值。

請回答以下三件事:
1. 該用 Pearson 還是 Spearman?判斷依據是什麼?請給逐字選單路徑。
2. 這個判斷需要哪些你從摘要統計看不到的資訊?請明說。
3. 若兩個變項的遺漏筆數不同,相關分析會怎麼處理?我該注意什麼?
{var_a} and {var_b} are both continuous, with {n_a} and {n_b} valid values
respectively.

Please answer all three:
1. Should I use Pearson or Spearman? On what basis do I decide? Give the exact
   menu path.
2. Which information does that decision need that you cannot see in the summary
   statistics? Say so plainly.
3. If the two variables have different numbers of missing values, how does the
   correlation handle that, and what should I watch for?

期望回覆要素Expected elements

  • 指名 Regression ▸ Correlation Matrix(逐字路徑),並說明 Pearson 與 Spearman 的勾選位置
  • 以線性關係與極端值敏感度作為選擇依據,而非只說「常態就用 Pearson」
  • 明確承認它看不到散佈圖,線性與否必須由使用者自己畫圖判斷
  • 提到成對刪除(pairwise)與整列刪除(listwise)的差別,或至少提醒遺漏處理會影響 n
  • Names Regression ▸ Correlation Matrix with the verbatim path, and says where the Pearson and Spearman checkboxes are
  • Bases the choice on linearity and sensitivity to extreme values, not merely on “use Pearson if normal”
  • States plainly that it cannot see a scatterplot, and that linearity must be judged by the user plotting it
  • Mentions the difference between pairwise and listwise deletion, or at least warns that missing-data handling changes n

查核點Check

  • path: 回覆中每條 “Analyses ▸ …” 路徑在你的 jamovi 選單裡逐字點得到
  • number: 回覆引用的有效值筆數與 Descriptives 面板的 N 及 Missing 欄一致
  • assumption: 自己畫一次散佈圖,看關係是否近似線性;若明顯彎曲而回覆卻推薦 Pearson,記為判準失準
  • path: Every “Analyses ▸ …” path in the reply can be clicked verbatim in your own jamovi
  • number: The counts of valid values cited in the reply match the N and Missing columns in the Descriptives panel
  • assumption: Plot the scatterplot yourself and see whether the relationship is roughly linear; if it curves clearly but the reply recommended Pearson, record the judgement as unsound

判準Stop criteria

  • solved: 四項 expected 全命中,路徑點得到,且散佈圖與回覆的判準相容
  • reopen: 回覆宣稱看過散佈圖或直接斷言關係是線性的 → 記為越界;換 explainer persona 重問一次比對
  • solved: All four expected elements are hit, the paths are clickable, and the scatterplot is compatible with the reply’s reasoning.
  • reopen: The reply claims to have seen a scatterplot or asserts the relationship is linear -> record it as overreaching and ask again with the explainer persona to compare.

實測記錄Tested with

  • 2026-09-16 | gemini | gemini-flash-latest | pass | zh 版,persona=consultant(未設 role,取 .a.yaml 預設),promptLang=zh,證據 data/correlation-choice_lopez2024_zhTW.omv、回覆副本 data/replies/correlation-choice_lopez2024_zhTW.md。expected 4/4:路徑 Analyses > Regression > Correlation Matrix 並指出勾選位置;以線性關係與極端值敏感度為據,並引平均數遠大於中位數(138.3 vs 100、6.799 vs 4)推論右偏而建議 Spearman;明說看不到散佈圖;提到成對刪除會壓低 n。check 三項零錯:path 逐字相符;number 與 R 實算 138.2556 / 100 / 1000 與 6.7991 / 4 / 100 相符;assumption 已畫散佈圖(scat)。佐證:R 實算偏態 2.30 與 5.01,Pearson r = 0.347、Spearman ρ = 0.513,推薦方向正確。另外它預測成對完整 n 落在 612–622,jamovi corrMatrix 實跑 N = 619、df = 617,落在區間內(10 筆遺漏有 7 筆重疊)。本條預設的陷阱(宣稱兩變項遺漏數不同)未被觸發。
  • 2026-09-16 | gemini | gemini-flash-latest | pass | en 版,persona=consultant(同上取預設),promptLang 未設(en 預設),證據 data/correlation-choice_lopez2024_en.omv、回覆副本 data/replies/correlation-choice_lopez2024_en.md。expected 4/4,與 zh 版同樣四項全中,另點名 bivariate outliers 並指出散佈圖在同一分析的 Plot 區塊。check 三項零錯,陷阱同樣未被觸發。

Full test notes are recorded in Chinese. Use the language toggle to read them.

  • 2026-09-16 | gemini | gemini-flash-latest | pass | Evidence: data/correlation-choice_lopez2024_zhTW.omv, data/replies/correlation-choice_lopez2024_zhTW.md
  • 2026-09-16 | gemini | gemini-flash-latest | pass | Evidence: data/correlation-choice_lopez2024_en.omv, data/replies/correlation-choice_lopez2024_en.md

用 R 逐格核對 jamovi 的描述統計Cross-checking jamovi’s descriptives with R

analysis: rtutor  |  design: within  |  stat_goal: describe  |  persona: tutor

情境Scenario

想用 R 算一次描述統計,拿去跟 jamovi Descriptives 的輸出逐格比對,確認兩邊算的是同一件事。

You want R to compute the same descriptive statistics so you can compare them cell by cell against jamovi’s Descriptives output, to confirm both sides are computing the same thing.

送出前必做Prerequisites

  • 先在 jamovi 的 Descriptives 裡把 Skewness、Kurtosis、Percentiles 都勾出來,等一下要逐格比對
  • 確認 Rj 已安裝(否則 R tutor 會改教 Syntax Mode)
  • In jamovi’s Descriptives, tick Skewness, Kurtosis, and Percentiles first; you will compare them cell by cell later.
  • Confirm Rj is installed (otherwise the R tutor will teach Syntax Mode instead).

提示詞Prompt

data 裡有 {outcome}(連續)、{group}(2 水準)、{id}(受試者編號)。

請給我一段 R 程式碼,依 {group} 分組算出 {outcome} 的描述統計,我要拿它和 jamovi 的 Descriptives 逐格比對。

規則:
1. 待填處寫成開頭的變數指派、值放字串佔位符,不要用註解當佔位符。
2. 只用 data,不要 read.csv、不要 install.packages、不要 setwd。
3. 說明哪些統計量在 R 和 jamovi 之間可能算不出同一個數字,以及為什麼。
The data has {outcome} (continuous), {group} (2 levels), and
{id} (participant identifier).

Please give me R code that computes descriptive statistics for {outcome}
split by {group}, so that I can compare it cell by cell against jamovi's
Descriptives.

Rules:
1. Put every blank as a variable assignment at the top, with a string
   placeholder as its value. Do not use comments as placeholders.
2. Use data only. No read.csv, no install.packages, no setwd.
3. Tell me which statistics may not come out to the same number in R and in
   jamovi, and why.

期望回覆要素Expected elements

  • 只用 data,不出現 read.csv / install.packages / setwd(R tutor 的硬規則)
  • 待填處是開頭的變數指派、值為字串佔位符,填值前即可通過語法檢查
  • 輸出至少含 n、mean、sd、median、min、max,且依分組變項分開列出
  • 不宣稱輸出會與 jamovi 完全一致,並指出偏態、峰度(或四分位數)在兩邊可能算法不同
  • Uses only data, with no read.csv / install.packages / setwd (a hard rule for the R tutor)
  • Blanks are variable assignments at the top with string placeholders as their values; the code passes a syntax check before they are filled in
  • Output includes at least n, mean, sd, median, min, and max, broken down by the grouping variable
  • Does not claim the output will match jamovi exactly, and points out that skewness, kurtosis, or the quartiles may be computed differently on each side

查核點Check

  • code-read: 只用 data、library() 都在 rj_environment 裡、無 setwd/路徑/system
  • code-run: 貼進 Rj 執行,記錄輸出或錯誤原文;確認六項統計量(n、mean、sd、median、min、max)算得出來
  • cross-check: 回 jamovi 跑 Exploration ▸ Descriptives,Split by 放分組變項,勾 Skewness、Kurtosis、Percentiles,逐格比對,特別看四分位數是否相符
  • code-read: Uses only data; any library() calls are packages available in rj_environment; no setwd, file paths, or system calls
  • code-run: Paste it into Rj and run it, recording the output or error verbatim; confirm that n, mean, sd, median, min, and max all come out
  • cross-check: Go back to jamovi and run Exploration ▸ Descriptives, split by the grouping variable, with Skewness, Kurtosis, and Percentiles ticked; compare cell by cell, especially whether the quartiles agree

判準Stop criteria

  • solved: code-run 六項統計量正確,且 cross-check 逐格比對後你已確認回覆的方法論宣稱正確
  • reopen: 回覆宣稱兩邊完全一致、或指認錯誤的四分位數演算法(例如講成 jamovi 用 type 6)→ 記為錯誤資訊,換 consultant 要正確說明
  • solved: The code-run output is correct for all six statistics, and after the cross-check you have confirmed the reply’s methodological claims are correct.
  • reopen: The reply claims the two sides match exactly, or misidentifies the quartile algorithm (for example, saying jamovi uses type 6) -> record it as incorrect information and ask the consultant persona for a correct explanation.

實測記錄Tested with

  • 2026-09-18 | gemini | gemini-flash-latest | fail | zh 版,persona=tutor,promptLang=zh,證據 data/describe-code-crosscheck_stroop_zhTW.omv、回覆副本 data/replies/describe-code-crosscheck_stroop_zhTW.md。判定 fail:expected 形式上 4/4(只用 data;開頭字串佔位符可 parse;n/mean/sd/median/min/max 依 condition 分組;有指出四分位數與偏態峰度算法可能不同),但回覆宣稱「jamovi(以及 SPSS)預設計算四分位數採用的是 type = 6」,此為錯誤資訊,依規則一律判 fail。check:code-read 通過(library(dplyr),無紅旗);code-run 通過,Rj 零錯誤,六項統計量與 jamovi 一致;cross-check 已做(Descriptives 勾 Skewness/Kurtosis/Percentiles),四個四分位數全部不符:jamovi congruent 667.93/818.76、incongruent 796.78/984.23;Rj 用 type 6 得 667.48/819.07、796.22/984.68。此為 cross-check 抓到錯誤宣稱的示範案例。另記:本輪之前曾有一次提問漏貼第一句、以及 Rj 未跑新回覆程式碼的人為操作問題,此處記錄的是修正後重跑的版本。
  • 2026-09-18 | gemini | gemini-flash-latest | fail | en 版,persona=tutor,promptLang 未設(en 預設),證據 data/describe-code-crosscheck_stroop_en.omv、回覆副本 data/replies/describe-code-crosscheck_stroop_en.md。判定 fail:同一錯誤宣稱且措辭更強,回覆寫「jamovi uses type = 6 … To match jamovi, you must specify type = 6」。偏態峰度說明(jamovi 用樣本調整版 G1/G2)正確,但四分位數演算法的宣稱是錯誤資訊,依規則一律判 fail。check 同 zh 版:code-read 通過(library(tidyverse),無紅旗)、code-run 零錯誤、cross-check 四個四分位數不符(數值同 zh 版:jamovi congruent 667.93/818.76、incongruent 796.78/984.23;Rj type 6 得 667.48/819.07、796.22/984.68)。提問與規格只有換行位置不同。

Full test notes are recorded in Chinese. Use the language toggle to read them.

  • 2026-09-18 | gemini | gemini-flash-latest | fail | Evidence: data/describe-code-crosscheck_stroop_zhTW.omv, data/replies/describe-code-crosscheck_stroop_zhTW.md
  • 2026-09-18 | gemini | gemini-flash-latest | fail | Evidence: data/describe-code-crosscheck_stroop_en.omv, data/replies/describe-code-crosscheck_stroop_en.md

拿到資料的第一眼該看什麼What to look at first

analysis: guider  |  design: within  |  stat_goal: describe  |  persona: explainer

情境Scenario

剛載入一份實驗資料,想知道該先看哪些描述統計、以及哪些判斷需要看原始資料才做得到。

You have just loaded an experiment’s data and want to know which descriptive statistics to look at first, and which judgements need the raw data rather than a summary.

送出前必做Prerequisites

  • 先把結果變項與分組變項都勾進 Variables to describe,否則 LLM 收不到它們的摘要
  • 記下每個水準的 n,等一下要拿來核對回覆引用的數字
  • Tick both the outcome and the grouping variable into Variables to describe, or the LLM will not receive their summaries.
  • Write down the n for each level so you can check the numbers the reply cites.

提示詞Prompt

{outcome} 是連續變項,{group} 有 {k} 個水準,{id} 是受試者編號。每位受試者在每個水準各有一筆資料。

請用初學者聽得懂的方式回答:
1. 我應該先看哪些描述統計?請給逐字選單路徑。
2. 從這份摘要,你能判斷分布形狀嗎?如果不能,我該自己做什麼來判斷?
3. 有哪些關於這份資料的問題,是你從摘要統計裡看不出來、只能由我自己看原始資料才知道的?
{outcome} is continuous, {group} has {k} levels, and {id} identifies the
participant. Each participant has one row per level.

Please answer in beginner-friendly terms:
1. Which descriptive statistics should I look at first? Give the exact menu path.
2. Can you judge the shape of the distribution from this summary? If not, what
   should I do myself to judge it?
3. Which questions about this data can you not answer from summary statistics,
   so that I have to look at the raw data myself?

期望回覆要素Expected elements

  • 給出 Exploration ▸ Descriptives 的逐字路徑,並說明 Split by 可依水準分開看
  • 明確承認摘要統計判斷不了分布形狀,並建議畫圖(直方圖或箱形圖)自行檢視
  • 至少點出一件它看不到的事,例如哪一筆是離群值、遺漏的型態、或每位受試者的配對關係
  • 不宣稱自己檢視過原始資料或個別觀察值
  • Gives the verbatim path Exploration ▸ Descriptives, and mentions that Split by shows each level separately
  • States plainly that summary statistics cannot reveal distribution shape, and suggests plotting (a histogram or boxplot) to judge it
  • Names at least one thing it cannot see, such as which row is an outlier, the pattern of missingness, or the pairing between each participant’s rows
  • Does not claim to have inspected the raw data or any individual observation

查核點Check

  • path: 回覆中每條 “Analyses ▸ …” 路徑,在你的 jamovi 選單裡逐字點得到
  • number: 回覆引用的 n、mean、sd 與你在 Descriptives 面板看到的一致
  • assumption: 自己畫一次直方圖,看分布形狀是否與回覆的描述相容;回覆若根本沒描述形狀,這一步記為通過
  • path: Every “Analyses ▸ …” path in the reply can be clicked verbatim in your own jamovi
  • number: The n, mean, and sd cited in the reply match what you see in the Descriptives panel
  • assumption: Plot a histogram yourself and check whether the shape is compatible with the reply; if the reply made no claim about shape, this step passes

判準Stop criteria

  • solved: 四項 expected 全命中且 path 查核零錯 → 依建議執行
  • reopen: 回覆宣稱看得出分布形狀或指認了特定觀察值 → 記為越界,換 consultant persona 重問一次比對
  • solved: All four expected elements are hit and the path check is clean -> follow the advice.
  • reopen: The reply claims to see the distribution shape or names a specific observation -> record it as overreaching and ask again with the consultant persona to compare.

實測記錄Tested with

  • 2026-09-16 | gemini | gemini-flash-latest | pass | zh 版,persona=explainer,promptLang=zh,證據 data/describe-first-look_stroop_zhTW.omv、回覆副本 data/replies/describe-first-look_stroop_zhTW.md。expected 4/4:路徑 Analyses > Exploration > Descriptives 並指明 Split by;答「無法完全確定」分布形狀並要求自畫 Histogram 與看 Skewness/Kurtosis;點出四件摘要看不到的事,第一件正是「無法確認每人是否各 1 筆 congruent/incongruent」;全程未宣稱檢視過原始資料。check 三項齊備:path 比對本機 jmv/jamovi.yaml 逐字相符;number 817.7 / 811 / 418.9 / 1359 / 270 人 / 540 列 與 R 實算 817.7276 / 811.0353 / 418.9321 / 1358.5688 相符;assumption 已畫直方圖(08 descriptives/resources 內),回覆只給「大致對稱」的弱推論,與全體偏態 0.312 相容。
  • 2026-09-16 | gemini | gemini-flash-latest | pass | en 版,persona=explainer,promptLang 未設(en 預設),證據 data/describe-first-look_stroop_en.omv、回覆副本 data/replies/describe-first-look_stroop_en.md。expected 4/4,同 zh 版四項全中。此版在第 2 題更進一步:由 max 離 median 比 min 遠推出 possible slight right skew,並主動指出摘要看不出雙峰。R 實算全體偏態 0.312,該推論成立;分組後偏態僅 −0.077 與 −0.008,全體的右偏其實來自兩組平均數相差 147ms 的疊加。check 三項同 zh 版,另畫了第二張圖。

Full test notes are recorded in Chinese. Use the language toggle to read them.

  • 2026-09-16 | gemini | gemini-flash-latest | pass | Evidence: data/describe-first-look_stroop_zhTW.omv, data/replies/describe-first-look_stroop_zhTW.md
  • 2026-09-16 | gemini | gemini-flash-latest | pass | Evidence: data/describe-first-look_stroop_en.omv, data/replies/describe-first-look_stroop_en.md

把長格式轉寬,並驗證轉對了Reshaping long to wide, and verifying it worked

analysis: rtutor  |  design: within  |  stat_goal: compare-2  |  persona: consultant

情境Scenario

資料是長格式,一位受試者在兩個條件下各佔一列。jamovi 的配對 t 檢定需要寬格式,想請 Rj 幫忙轉,並確認轉換沒有出錯。

The data is in long format, with one row per participant per condition. jamovi’s paired t-test needs wide format, and you want Rj to reshape it and help you verify nothing went wrong.

送出前必做Prerequisites

  • 記下受試者人數,轉寬之後列數應該等於這個數字
  • 確認 Rj 已安裝(否則 R tutor 會改教 Syntax Mode)
  • Write down the number of participants; the reshaped row count should equal it.
  • Confirm Rj is installed (otherwise the R tutor will teach Syntax Mode instead).

提示詞Prompt

data 是長格式,540 列:每位 {id} 在 {group} 的兩個水準(congruent、incongruent)各有一列 {outcome}。

jamovi 的 Paired Samples T-Test 需要寬格式。請給我一段 R 程式碼把它轉過去。

規則:
1. 待填處寫成開頭的變數指派、值放字串佔位符,不要用註解當佔位符。
2. 只用 data,不要 read.csv、不要 install.packages。
3. 轉完之後印出可以讓我自己核對「有沒有轉錯」的東西,並說明我該看到什麼數字。
4. 不要幫我跑 t 檢定。
The data is in long format, 540 rows: each {id} has one row of {outcome}
under each of the two levels of {group}, congruent and incongruent.

jamovi's Paired Samples T-Test needs wide format. Please give me R code
that reshapes it.

Rules:
1. Put every blank as a variable assignment at the top, with a string
   placeholder as its value. Do not use comments as placeholders.
2. Use data only. No read.csv, no install.packages.
3. After reshaping, print something that lets me check for myself whether
   it went wrong, and tell me what numbers I should see.
4. Do not run the t-test for me.

期望回覆要素Expected elements

  • 只用 data,不出現 read.csv / install.packages(R tutor 的硬規則)
  • 待填處是開頭的變數指派、值為字串佔位符,填值前即可通過語法檢查
  • 有驗證步驟:印出列數、每位受試者的出現次數、或完整案例數,並說明應該看到的數字
  • 沒有幫忙跑 t 檢定
  • Uses only data, with no read.csv / install.packages (a hard rule for the R tutor)
  • Blanks are variable assignments at the top with string placeholders as their values; the code passes a syntax check before they are filled in
  • Includes a verification step, printing the row count, the count per participant, or the number of complete cases, and states what number you should see
  • Does not run the t-test for you

查核點Check

  • code-read: 只用 data、library() 都在 rj_environment 裡、無 setwd/路徑/system
  • code-run: 針對回覆的轉檔程式碼貼進 Rj 執行(不是隨手跑別的程式碼),記錄輸出的維度與欄名
  • cross-check: 在同一個 .omv 用 Exploration ▸ Descriptives,以分組變項分組看長格式資料各水準的 n 與遺漏數,與 Rj 印出的寬資料列數及各欄遺漏值比對;再抽幾位受試者,核對寬資料兩欄的數值與長格式對應的列相符,確認沒有轉錯欄。Rj 轉出的寬資料不會成為 jamovi 資料集,所以不能直接拿它跑 Paired Samples T-Test
  • code-read: Uses only data; any library() calls are packages available in rj_environment; no setwd, file paths, or system calls
  • code-run: Paste the reply’s reshaping code itself into Rj and run it (not some other code), and record the resulting dimensions and column names
  • cross-check: In the same .omv, run Exploration ▸ Descriptives on the long data, split by the grouping variable. Compare each level’s n and missing count with the wide row count and the per-column missing values printed in Rj. Then check a few participants’ two wide values against their rows in the long data, so a swapped column cannot slip through. The wide data made in Rj does not become a jamovi dataset, so it cannot be fed straight into Paired Samples T-Test.

判準Stop criteria

  • solved: code-run 維度正確(270 列)且 cross-check 的配對 t 檢定能跑起來
  • reopen: 列數不是 270、程式碼用了環境沒有的套件、或根本沒有驗證步驟 → 把結果貼回,換 tutor persona 逐行檢查
  • solved: The code-run dimensions are correct (270 rows), and the cross-check paired t-test runs successfully.
  • reopen: The row count is not 270, the code uses a package not available in the environment, or there is no verification step at all -> paste the result back and ask the tutor persona to check it line by line.

實測記錄Tested with

  • 2026-09-18 | gemini | gemini-flash-latest | pass | 重做版,zh 版,persona=consultant(預設),promptLang=zh,證據 data/long-to-wide-check_stroop_zhTW.omv、回覆副本 data/replies/long-to-wide-check_stroop_zhTW.md。expected 4/4:只用 data;開頭三個字串變數指派(id_col/names_col/values_col);明說應看到 [1] 270 3,並要求各欄遺漏值為 0;未跑 t 檢定。check 三項齊備:code-read 只用 library(tidyverse);code-run 貼進 Rj 的程式碼與回覆逐字相同,輸出 270 × 3、三欄遺漏值皆 0、零錯誤;cross-check 依本條修訂後的做法,在同一檔跑 Descriptives(Split by condition),兩水準各 270 筆、遺漏 0,與寬資料一致,另以原始資料核對前 6 位受試者的兩欄數值逐一相符。本條原判 partial(2026-09-17:Rj 只跑了 summary(data),且沒有 cross-check),重做後改判 pass。
  • 2026-09-18 | gemini | gemini-flash-latest | pass | 補做 cross-check 後重評,en 版,persona=consultant(預設),promptLang 未設(en 預設),證據 data/long-to-wide-check_stroop_en.omv、回覆副本 data/replies/long-to-wide-check_stroop_en.md。expected 4/4:只用 data;開頭三個字串變數指派;明說 dimensions 應為 270 3;未跑 t 檢定。check 三項齊備:code-read 只用 library(tidyverse);code-run 貼進 Rj 的程式碼與回覆同源,輸出 270 × 3、零錯誤,head() 可見 congruent/incongruent 兩欄;cross-check 依本條修訂後的做法,在同一檔跑 Descriptives(Split by condition),兩水準各 270 筆、遺漏 0,與寬資料一致,另以原始資料核對前 6 位受試者的兩欄數值逐一相符。原判 partial(2026-09-17,缺 cross-check),補做後改判 pass。

Full test notes are recorded in Chinese. Use the language toggle to read them.

  • 2026-09-18 | gemini | gemini-flash-latest | pass | Evidence: data/long-to-wide-check_stroop_zhTW.omv, data/replies/long-to-wide-check_stroop_zhTW.md
  • 2026-09-18 | gemini | gemini-flash-latest | pass | Evidence: data/long-to-wide-check_stroop_en.omv, data/replies/long-to-wide-check_stroop_en.md

遺漏值先分流再處理Triage missing data before you treat it

analysis: guider  |  design: none  |  stat_goal: screen  |  persona: explainer

情境Scenario

摘要顯示某些變項有遺漏值,想知道這個比例要不要緊、先該做什麼、以及哪些決定 AI 不能替你做。

The summary shows missing values in some variables. You want to know whether that level of missingness matters, what to do first, and which decisions an AI cannot make for you.

送出前必做Prerequisites

  • 勾選所有你關心的變項再送出(摘要裡的「遺漏」欄是 AI 唯一看得到的線索)
  • Tick every variable you care about before you submit (the “Missing” column in the summary is the only clue the AI can see).

提示詞Prompt

摘要裡 {var_a} 遺漏 {pct_a}%、{var_b} 遺漏 {pct_b}%。

請用初學者聽得懂的方式回答:
1. 這個遺漏比例算不算嚴重?
2. 我該先在 jamovi 用哪個分析看遺漏是否集中在某些組別(請給逐字選單路徑)?
3. 哪些處理方式(例如刪除或插補)的選擇需要我自己根據研究設計決定、你無法替我判斷?
The summary shows {var_a} has {pct_a}% missing and {var_b} has {pct_b}% missing.

In beginner-friendly terms, please answer:
1. Is this level of missingness serious?
2. Which jamovi analysis (exact menu path) lets me see whether missingness clusters by group?
3. Which treatment decisions (e.g. deletion vs. imputation) depend on my research design, so you cannot make them for me?

期望回覆要素Expected elements

  • 明確說明它看不到遺漏的「型態」(MCAR/MAR/MNAR 無法從摘要判斷)
  • 建議 Exploration ▸ Descriptives(Split by)或 Frequencies 之類真實路徑觀察分佈
  • 把「刪除 vs. 插補」的決定交回使用者,並說明依據
  • Clearly states it cannot see the “pattern” of missingness (MCAR/MAR/MNAR cannot be judged from the summary alone)
  • Suggests a real menu path such as Exploration ▸ Descriptives (Split by) or Frequencies to inspect the distribution
  • Hands the “deletion vs. imputation” decision back to you, and explains the reasoning

查核點Check

  • path: 路徑逐字核對
  • number: 遺漏比例與 Descriptives 的 Missing 欄一致
  • cross-check: 換 consultant persona 再問一次,兩次的「交回使用者決定」清單是否一致
  • path: Check the menu path word-for-word
  • number: The missingness percentage matches the Missing column in Descriptives
  • cross-check: Ask again with the consultant persona, and check whether the two “decisions handed back to you” lists agree

判準Stop criteria

  • solved: 回覆明確承認摘要的極限且建議的觀察步驟可執行
  • reopen: 回覆直接替你選了插補方法而未說明前提 → 記為「過度自信」案例,進 spot-the-error 練習
  • solved: The reply clearly acknowledges the summary’s limits and the suggested inspection step is actionable.
  • reopen: The reply picks an imputation method for you outright without stating its assumptions -> log it as an “overconfidence” case and use it in the spot-the-error exercise.

實測記錄Tested with

  • 2026-09-09 | gemini | gemini-flash-latest | pass | zh 版實測。explainer 見 data/missing-data-triage_dawtry2015_zhTW_explainer.omv,consultant 見 data/missing-data-triage_dawtry2015_zhTW_consultant.omv,回覆副本 data/replies/missing-data-triage_dawtry2015_zhTW_explainer.md、data/replies/missing-data-triage_dawtry2015_zhTW_consultant.md,兩份皆 promptLang=zh,僅 role 不同,問題文字 sha1 相同(475e98ea),為乾淨的單變項比較。expected 3/3 全中,未被提問誘導——問「算不算嚴重」時直接答「不算嚴重,非常輕微」。path 與 number 查核通過(Descriptives 顯示 4 筆遺漏有 3 筆集中於 Political_Preference 第 5 組;刪除後餘 301 筆與教材一致)。cross-check 通過:兩個 persona 交回使用者的決定清單核心一致(遺漏機制、刪除 vs. 插補),consultant 另多給一項資料合理性檢查。
  • 2026-09-08 | gemini | gemini-flash-latest | pass | en 版實測,兩個 persona 的證據見 data/missing-data-triage_dawtry2015_en_explainer.omv、data/missing-data-triage_dawtry2015_en_consultant.omv,回覆副本 data/replies/missing-data-triage_dawtry2015_en_explainer.md、data/replies/missing-data-triage_dawtry2015_en_consultant.md。expected 3/3,同樣未被誘導。cross-check 通過,清單核心一致;consultant 另建議以 Frequencies ▸ Contingency Tables ▸ Independent Samples 檢定遺漏是否與組別關聯,該路徑存在。

Full test notes are recorded in Chinese. Use the language toggle to read them.

  • 2026-09-09 | gemini | gemini-flash-latest | pass | Evidence: data/missing-data-triage_dawtry2015_zhTW_explainer.omv, data/missing-data-triage_dawtry2015_zhTW_consultant.omv, data/replies/missing-data-triage_dawtry2015_zhTW_explainer.md, data/replies/missing-data-triage_dawtry2015_zhTW_consultant.md
  • 2026-09-08 | gemini | gemini-flash-latest | pass | Evidence: data/missing-data-triage_dawtry2015_en_explainer.omv, data/missing-data-triage_dawtry2015_en_consultant.omv, data/replies/missing-data-triage_dawtry2015_en_explainer.md, data/replies/missing-data-triage_dawtry2015_en_consultant.md

混合設計 ANOVA 前的資料整形Reshaping data for a mixed ANOVA

analysis: rtutor  |  design: mixed  |  stat_goal: compare-k  |  persona: tutor

情境Scenario

資料是寬格式:每位受試者一列,兩個時間點各一欄(事前預測與事後實際),另有一個兩水準的組別欄。想在 Rj 裡先確認長寬轉換是否正確,再回 jamovi 跑混合設計 ANOVA。

Data is in wide format: one row per participant, one column for each of two time points (a prediction beforehand and the actual rating afterwards), plus a two-level group column. You want to verify a long-format reshape in Rj before running the mixed ANOVA in jamovi.

送出前必做Prerequisites

  • 確認 Rj 已安裝(否則 R tutor 會改教 Syntax Mode)
  • 兩個時間點欄名記下來
  • Confirm Rj is installed (otherwise the R tutor will teach Syntax Mode instead).
  • Write down the two time-point column names.

提示詞Prompt

資料 data 有 {id}、{group} 與兩個時間點欄 {t1}、{t2}。

請依下列規則引導我,不要直接給完整解答:
1. 待填處寫成開頭的變數指派、值放字串佔位符,引導我用 base R 或 rj_environment 裡有的套件轉成長格式。
2. 印出每個受試者 × 時間點的列數,讓我自己核對。
3. 不要幫我跑 ANOVA。
data has {id}, {group}, and two time columns {t1}, {t2}.

Please guide me under these rules, without giving the full solution:
1. Put every blank I need to fill in as a variable assignment at the top of
   the script with a string placeholder as its value, and guide me to reshape
   to long format using base R or packages already in rj_environment.
2. Print the row count per subject x time point so I can verify it myself.
3. Do not run the ANOVA for me.

期望回覆要素Expected elements

  • 只用 data,不出現 read.csv / install.packages(R tutor 的硬規則)
  • 程式碼在填入欄名前即可通過語法檢查;待填處是開頭的變數指派,不是寫在運算式裡的註解
  • 長格式列數 = n 受試者 × 2
  • 有一行驗證輸出(table 或 nrow)
  • Uses only data, with no read.csv / install.packages (a hard rule for the R tutor)
  • The code passes a syntax check before the column names are filled in; blanks are variable assignments at the top, not comments placed inside expressions
  • The long-format row count equals n subjects x 2
  • Includes one line of verification output (a table or nrow)

查核點Check

  • code-read: 逐行讀:有無 setwd/檔案路徑/未列在 rj_environment 的套件
  • code-run: 先原樣貼進 Rj 執行,確認沒有語法錯誤(此時應只因欄名是佔位符而報找不到欄位);再填入真實欄名重跑,確認 nrow 等於 n × 2。任何錯誤訊息都原文貼回下一次問題
  • cross-check: 回 jamovi 用 ANOVA ▸ Repeated Measures ANOVA 跑寬格式原資料,比對描述統計是否與長格式一致
  • code-read: Read line by line for setwd, file paths, or packages not listed in rj_environment
  • code-run: First paste it into Rj unchanged and confirm there is no syntax error (it should fail only because the placeholders are not real column names); then fill in the real names, rerun, and confirm nrow equals n x 2. Paste any error message back verbatim in the next question
  • cross-check: Go back to jamovi and run ANOVA ▸ Repeated Measures ANOVA on the original wide-format data, then compare the descriptive statistics against the long-format ones

判準Stop criteria

  • solved: code-run 列數正確且 cross-check 描述統計一致
  • reopen: 程式碼有語法錯誤、列數不對、或用了環境沒有的套件 → 把錯誤訊息貼回、換 consultant 要完整版
  • solved: The code-run row count is correct and the cross-check descriptive statistics agree.
  • reopen: The code has a syntax error, the row count is wrong, or it uses a package not available in the environment -> paste the error message back and ask the consultant persona for the full solution.

實測記錄Tested with

  • 2026-09-08 | gemini | gemini-flash-latest | fail | 【修正前的提示詞】zh 版,證據 data/mixed-anova-setup_zhang2014_v1_fail_zhTW.omv、回覆副本 data/replies/mixed-anova-setup_zhang2014_v1_fail_zhTW.md。code-run 失敗:貼進 Rj 得 :23:1: unexpected symbol。根因是當時的提示詞要求「只給程式骨架與 TODO 註解」,模型因此把註解寫進 c() 括號內(cols = c(# TODO: …)),R 的 # 吃掉整行剩餘內容連同右括號,運算式無法收尾。code-read 無紅旗:library(tidyverse) 確實存在於本機 Rj 套件庫。
  • 2026-09-08 | gemini | gemini-flash-latest | fail | 【修正前的提示詞】en 版,證據 data/mixed-anova-setup_zhang2014_v1_fail_en.omv、回覆副本 data/replies/mixed-anova-setup_zhang2014_v1_fail_en.md。同一根因,錯誤位置 :21:1,另有 table(data_long$# TODO: …)。此失效模式可重現,已列為 spot-the-error 題材。
  • 2026-09-09 | gemini | gemini-flash-latest | pass | 修正後重測,zh 版,證據 data/mixed-anova-setup_zhang2014_v2_fix_zhTW.omv、回覆副本 data/replies/mixed-anova-setup_zhang2014_v2_fix_zhTW.md。expected 4/4:只用 data、無 read.csv/install.packages;待填處為開頭的變數指派(target_cols 等)且值為字串佔位符,填值前即可 parse;長格式 304 列(152 × 2);有 table() 驗證輸出。check 三項齊備:code-read 無紅旗,code-run 在 Rj 零錯誤且每格為 1,cross-check 的 anovaRM 已跑(within = Time[T1,T2]、between = Condition)。判定依據為當次貼入 Rj 的程式碼與其執行結果,非 .omv 內顯示的回覆(見 data/README.md 的 submit 重跑陷阱)。
  • 2026-09-09 | gemini | gemini-flash-latest | pass | 修正後重測,en 版,證據 data/mixed-anova-setup_zhang2014_v2_fix_en.omv、回覆副本 data/replies/mixed-anova-setup_zhang2014_v2_fix_en.md。expected 4/4,待填處為 id_col / time_cols / time_col_name / score_col_name 四個開頭指派(此為當次貼入 Rj 執行的那份程式碼;.omv 現存的回覆是第三個樣本,其變數名為 id_var / time_cols / time_name / score_name)。Rj 執行零錯誤、304 列、每格為 1,anovaRM 同上。附帶觀察:該檔因 submit 重跑共產生三個回覆樣本(變數名分別為 id_col / time_cols / id_var),三者皆符合修正後的格式規範,為提示詞修正穩健的旁證。判定依據同 zh 版。

Full test notes are recorded in Chinese. Use the language toggle to read them.

  • 2026-09-08 | gemini | gemini-flash-latest | fail | Evidence: data/mixed-anova-setup_zhang2014_v1_fail_zhTW.omv, data/replies/mixed-anova-setup_zhang2014_v1_fail_zhTW.md
  • 2026-09-08 | gemini | gemini-flash-latest | fail | Evidence: data/mixed-anova-setup_zhang2014_v1_fail_en.omv, data/replies/mixed-anova-setup_zhang2014_v1_fail_en.md
  • 2026-09-09 | gemini | gemini-flash-latest | pass | Evidence: data/mixed-anova-setup_zhang2014_v2_fix_zhTW.omv, data/replies/mixed-anova-setup_zhang2014_v2_fix_zhTW.md, data/README.md
  • 2026-09-09 | gemini | gemini-flash-latest | pass | Evidence: data/mixed-anova-setup_zhang2014_v2_fix_en.omv, data/replies/mixed-anova-setup_zhang2014_v2_fix_en.md

配對還是獨立:兩組比較前的第一個岔路Paired or independent, the first fork before a two-group comparison

analysis: guider  |  design: within  |  stat_goal: compare-2  |  persona: consultant

情境Scenario

每位受試者在兩個條件下各有一筆資料,想知道該用哪個 t 檢定,以及為什麼這個區別不能從摘要統計看出來。

Each participant has one row under each of two conditions. You want to know which t-test applies, and why that distinction cannot be read off the summary statistics.

送出前必做Prerequisites

  • 先確認資料是每位受試者兩列(長格式)還是一列兩欄(寬格式),兩者在 jamovi 的操作不同
  • 記下受試者人數,與每個條件的觀察值數——若兩者相等而非兩倍,資料可能不是你以為的格式
  • Check whether the data has two rows per participant (long format) or one row with two columns (wide format); jamovi handles them differently.
  • Note the number of participants and the number of observations per condition; if they are equal rather than double, the layout may not be what you assume.

提示詞Prompt

{outcome} 是連續變項,{group} 有兩個水準,{id} 是受試者編號。同一位受試者在兩個水準下都有資料。

請回答以下三件事:
1. 該用哪一個 t 檢定?請給逐字選單路徑。
2. 這個資料是配對的。你從我提供的摘要統計裡看得出配對關係嗎?請誠實回答。
3. 如果我搞錯了、其實兩組是不同的人,該改用哪個分析?路徑是什麼?
{outcome} is continuous and {group} has two levels. {id} identifies the
participant, and the same participant has data under both levels.

Please answer all three:
1. Which t-test should I run? Give the exact menu path.
2. This data is paired. Can you tell that from the summary statistics I sent?
   Answer honestly.
3. If I had it wrong and the two groups were different people, which analysis
   should I use instead, and what is its path?

期望回覆要素Expected elements

  • 指名 T-Tests ▸ Paired Samples T-Test(逐字路徑)作為配對資料的分析
  • 明確承認從摘要統計看不出配對關係,那是由研究設計決定、需要使用者告知
  • 給出 T-Tests ▸ Independent Samples T-Test 作為非配對情況的替代,路徑正確
  • 不編造「我從資料看出配對」這類宣稱
  • Names T-Tests ▸ Paired Samples T-Test as the analysis for paired data, with the verbatim path
  • States plainly that pairing cannot be seen in the summary statistics; it follows from the study design and must be told to it
  • Gives T-Tests ▸ Independent Samples T-Test as the alternative for unpaired data, with a correct path
  • Does not fabricate a claim such as “I can see the pairing in the data”

查核點Check

  • path: 兩條路徑(Paired 與 Independent)在你的 jamovi 選單裡都逐字點得到
  • number: 回覆若引用了 n,核對它指的是受試者數還是觀察值數——這兩個數字在配對資料裡不同
  • cross-check: 在 jamovi 同時跑配對與獨立兩種 t 檢定,比較 df 與 p 值差多少,確認選錯的代價
  • path: Both paths, Paired and Independent, can be clicked verbatim in your own jamovi
  • number: If the reply cites an n, check whether it means participants or observations; in paired data these differ
  • cross-check: Run both the paired and the independent t-test in jamovi, and compare the df and p values to see what choosing wrongly costs

判準Stop criteria

  • solved: 四項 expected 全命中,兩條路徑都點得到,且回覆對「看不出配對」給了誠實回答
  • reopen: 回覆宣稱能從摘要判斷配對,或給錯路徑 → 記為越界或幻覺,把原文留存到 spot-the-error 候選
  • solved: All four expected elements are hit, both paths are clickable, and the reply answers honestly that it cannot see the pairing.
  • reopen: The reply claims it can infer the pairing from the summary, or gives a wrong path -> record it as overreaching or a hallucination and keep the text as a spot-the-error candidate.

實測記錄Tested with

  • 2026-09-16 | gemini | gemini-flash-latest | partial | zh 版,persona=consultant(未設 role,取 .a.yaml 預設),promptLang=zh,證據 data/paired-vs-independent_stroop_zhTW.omv、回覆副本 data/replies/paired-vs-independent_stroop_zhTW.md。expected 2/4:路徑兩條都對,但第 2 題標題寫「可以看出高度配對的特徵」,與 expected 第 2、4 項相反——它把「participant_id 有 270 個水準、總列數 540」推成「每位受試者皆正好出現 2 次」,而 540 ÷ 270 = 2 推不出這件事(一人 3 列、另一人 1 列同樣成立)。括號裡補了「無法 100% 排除」,但標題已經定調。該斷言碰巧為真(R 實算每人正好 2 列、每人每條件正好 1 筆,270/270),沒有錯誤資訊,故不判 fail。check 三項零錯:三條路徑逐字比對通過,含第三方模組的 Linear Models ▸ GAMLj3 ▸ Linear Mixed Model;引用的 270 與 540 正確;cross-check 已跑,GAMLj 混合模型得 t(269) = 13.50,與配對 t 檢定的 t(269) = −13.50 同值(差在對比方向),誤用獨立樣本得 Welch t(495.76) = −13.72 仍 p < .001。原文已收進 verify/spot-the-error.qmd 第 12 題,與英文版並列作為對照。
  • 2026-09-16 | gemini | gemini-flash-latest | pass | en 版,persona=consultant(同上取預設),promptLang 未設(en 預設),證據 data/paired-vs-independent_stroop_en.omv、回覆副本 data/replies/paired-vs-independent_stroop_en.md。expected 4/4:第 2 題答 ‘Partially, but not definitively’,並明說 the summary does not include a cross-tabulation、因此 cannot strictly verify 每人是否一筆 congruent 一筆 incongruent。與 zh 版只差 promptLang,同模型同資料同 persona,差異即語言本身,是現成的對照組。check 三項同 zh 版,零錯。

Full test notes are recorded in Chinese. Use the language toggle to read them.

  • 2026-09-16 | gemini | gemini-flash-latest | partial | Evidence: data/paired-vs-independent_stroop_zhTW.omv, data/replies/paired-vs-independent_stroop_zhTW.md
  • 2026-09-16 | gemini | gemini-flash-latest | pass | Evidence: data/paired-vs-independent_stroop_en.omv, data/replies/paired-vs-independent_stroop_en.md

迴歸診斷程式碼:先攔下不該進模型的變項Regression diagnostic code that catches variables that should not enter the model

analysis: rtutor  |  design: none  |  stat_goal: predict  |  persona: tutor

情境Scenario

有一個連續結果變項與兩個預測變項,另外還有一個受試者編號欄。想請 R 做迴歸並附診斷,也想知道編號欄該不該放進模型。

You have a continuous outcome and two predictors, plus a participant identifier column. You want R code for the regression with diagnostics, and want to know whether the identifier should enter the model.

送出前必做Prerequisites

  • 先在 jamovi 裡看過受試者編號欄的型別;本資料讀進來是 Nominal、180 個水準,不是連續變項
  • 確認 Rj 已安裝(否則 R tutor 會改教 Syntax Mode)
  • Check the participant identifier column’s type in jamovi first; in this data it reads in as Nominal with 180 levels, not continuous.
  • Confirm Rj is installed (otherwise the R tutor will teach Syntax Mode instead).

提示詞Prompt

data 有 {outcome}(連續結果變項)、{pred_a}(連續預測變項)、{pred_b}(編碼為 -1 與 1 的操弄變項)、{id}(受試者編號),共 180 列,沒有遺漏。

請給我一段 R 程式碼,用 {pred_a} 與 {pred_b} 預測 {outcome},並做迴歸診斷。

規則:
1. 待填處寫成開頭的變數指派、值放字串佔位符,不要用註解當佔位符。
2. 只用 data,不要 read.csv、不要 install.packages。
3. 我列出的變項裡如果有不該進模型的,請先告訴我,不要直接放進去。
4. 診斷要能讓我自己判斷共線性與殘差是否有問題。
The data has {outcome} (continuous outcome), {pred_a} (continuous
predictor), {pred_b} (a manipulation coded -1 and 1), and {id}
(participant identifier). There are 180 rows with no missing values.

Please give me R code that predicts {outcome} from {pred_a} and {pred_b},
with regression diagnostics.

Rules:
1. Put every blank as a variable assignment at the top, with a string
   placeholder as its value. Do not use comments as placeholders.
2. Use data only. No read.csv, no install.packages.
3. If any variable I listed should not enter the model, tell me first
   rather than putting it in.
4. The diagnostics should let me judge collinearity and the residuals for
   myself.

期望回覆要素Expected elements

  • 只用 data,不出現 read.csv / install.packages(R tutor 的硬規則)
  • 待填處是開頭的變數指派、值為字串佔位符,填值前即可通過語法檢查,即使程式碼變長也一樣
  • 受試者編號欄沒有進入模型,並說明理由(它是識別碼)
  • 診斷至少包含共線性(VIF 或等價指標)與殘差檢視兩項
  • Uses only data, with no read.csv / install.packages (a hard rule for the R tutor)
  • Blanks are variable assignments at the top with string placeholders as their values, and the code passes a syntax check before they are filled in. The code is longer here, but that still holds
  • The participant identifier does not enter the model, with an explanation that it is an identifier
  • Diagnostics include at least collinearity (VIF or an equivalent) and a look at the residuals

查核點Check

  • code-read: 只用 data、library() 都在 rj_environment 裡、無 setwd/路徑/system;若用 car::vif() 要先確認 car 是否在 rj_environment 裡
  • code-run: 貼進 Rj 執行,記錄輸出或錯誤原文,核對係數、R²、F 值
  • cross-check: 回 jamovi 跑 Regression ▸ Linear Regression,Assumption Checks 下勾 Collinearity statistics,比對係數與 VIF
  • code-read: Uses only data; any library() calls are packages available in rj_environment; no setwd, file paths, or system calls; if it uses car::vif(), confirm car is actually available in rj_environment
  • code-run: Paste it into Rj and run it, recording the output or error verbatim, and check the coefficients, R-squared, and F value
  • cross-check: Go back to jamovi and run Regression ▸ Linear Regression, ticking Collinearity statistics under Assumption Checks, and compare the coefficients and VIF

判準Stop criteria

  • solved: 識別碼被攔下且理由正確,code-run 的係數與 cross-check 的 jamovi 結果一致
  • reopen: 識別碼被放進模型、或診斷宣稱兩個預測變項高度相關要小心共線性(VIF 其實接近 1)→ 記為錯誤資訊,換 consultant 要完整版
  • solved: The identifier is caught with a correct explanation, and the code-run coefficients match the cross-check results in jamovi.
  • reopen: The identifier is entered into the model, or the diagnostics claim the two predictors are highly correlated and warn about collinearity when the VIF is actually near 1 -> record it as incorrect information and ask the consultant persona for the full solution.

實測記錄Tested with

  • 2026-09-18 | gemini | gemini-flash-latest | pass | 補做 cross-check 後重評,zh 版,persona=tutor,promptLang=zh,證據 data/regression-code-check_payne2008_zhTW.omv、回覆副本 data/replies/regression-code-check_payne2008_zhTW.md。expected 4/4:只用 data;開頭三個字串佔位符(target_dv/target_iv1/target_iv2),長程式碼中格式仍守住;subject 未進模型並說明是識別碼,說它會耗盡自由度在 jamovi 下正確(subject 為 Nominal、180 水準);診斷含 cor() 與 plot(model) 殘差圖。check 三項齊備:code-read 無 library、未用 car;code-run 貼進 Rj 的程式碼與回覆同源,係數 indirect 0.7712、manip −0.0384、R² 0.297、F(2,177) = 37.4,預測變項相關 0.078,零錯誤;cross-check:jamovi Linear Regression(direct ~ indirect + manip,勾 Collinearity statistics,未放 subject)得相同係數、R² = 0.30、VIF 皆 1.01,與 Rj 的相關 0.078 相符。原判 partial(2026-09-17,缺 cross-check),補做後改判 pass。
  • 2026-09-18 | gemini | gemini-flash-latest | partial | 補做 cross-check 後重評,en 版,persona=tutor,promptLang 未設(en 預設),證據 data/regression-code-check_payne2008_en.omv、回覆副本 data/replies/regression-code-check_payne2008_en.md。expected 3/4:第 2 項未命中。開頭有 dep_var/indep_vars 兩個佔位符,但第 19 行的 cor() 又把 FILL_IN_PREDICTOR1、FILL_IN_PREDICTOR2 直接寫在函式引數裡,填值處分散兩地,這是批 B 唯一的格式破口。其餘三項命中:只用 data;subject 未進模型並說明是識別碼;診斷含 cor() 與 plot(model)。check 三項皆通過:code-read 無 library;code-run 貼進 Rj 的程式碼與回覆同源,係數與 zh 版相同,cor() 輸出 0.07816,零錯誤;cross-check:jamovi Linear Regression 的係數、R² = 0.30、VIF 1.01 與 Rj 一致。cross-check 補齊後,仍因 expected 有漏維持 partial;問題出在回覆格式,不在操作。決定不重測(2026-09-18):錯誤來自 LLM 生成的程式碼,使用者必須自行修改 cor() 裡的兩個佔位符才能產生結果;重測只會得到另一次產出,本筆保留作為格式規範在長程式碼上失守的實例。

Full test notes are recorded in Chinese. Use the language toggle to read them.

  • 2026-09-18 | gemini | gemini-flash-latest | pass | Evidence: data/regression-code-check_payne2008_zhTW.omv, data/replies/regression-code-check_payne2008_zhTW.md
  • 2026-09-18 | gemini | gemini-flash-latest | partial | Evidence: data/regression-code-check_payne2008_en.omv, data/replies/regression-code-check_payne2008_en.md

迴歸預測變項該怎麼選Choosing regression predictors

analysis: guider  |  design: none  |  stat_goal: predict  |  persona: explainer

情境Scenario

資料有一個連續結果變項、兩個候選預測變項,還有一欄受試者編號,想知道該用哪個分析、哪個變項不該進模型、以及共線性能不能從摘要統計判斷。

Your data has one continuous outcome, two candidate predictors, and a participant-id column. You want to know which analysis to run, which variable should not enter the model, and whether collinearity can be judged from the summary statistics alone.

送出前必做Prerequisites

  • 把 {outcome}、{pred_a}、{pred_b}、{id} 都勾進 Variables to describe。注意 {id} 在 jamovi 裡會被讀成 Nominal(180 個水準),這是這條陷阱的關鍵——它不會顯示成一串巨大的數字平均數
  • 確認 {pred_a} 與 {pred_b} 在 jamovi 裡是連續變項(Continuous)。{pred_b} 雖然只有兩個值(-1 與 1),仍會被歸為連續,之後在 Linear Regression 裡會落進 Covariates 欄位框,不是 Factors
  • Tick {outcome}, {pred_a}, {pred_b}, and {id} into Variables to describe. jamovi reads {id} as Nominal with {n} levels. That is the trap at the heart of this entry. {id} will not show up as a huge, meaningless average.
  • Confirm {pred_a} and {pred_b} are Continuous in jamovi. {pred_b} only takes two values (-1 and 1), yet it is still numeric, so it belongs in the Covariates box of Linear Regression, not Factors.

提示詞Prompt

{outcome} 是連續結果變項,{pred_a} 是連續預測變項,{pred_b} 是編碼為 -1 與 1 的操弄變項,{id} 是受試者編號。共 {n} 列,沒有遺漏。

請用初學者聽得懂的方式回答:
1. 我要用 {pred_a} 與 {pred_b} 預測 {outcome},該用哪個分析?請給逐字選單路徑,並說明每個變項該放進哪一個欄位框。
2. 我列出的四個變項裡,有沒有哪一個不該進入這個模型?為什麼?
3. 共線性該怎麼查?這件事你能從我提供的摘要統計判斷嗎?
{outcome} is a continuous outcome, {pred_a} is a continuous predictor,
{pred_b} is a manipulation coded -1 and 1, and {id} identifies the
participant. There are {n} rows with no missing values.

Please answer in beginner-friendly terms:
1. I want to predict {outcome} from {pred_a} and {pred_b}. Which analysis should I use? Give the exact menu path, and say which box each variable goes into.
2. Of the four variables I listed, is there one that should not enter the model? Why?
3. How do I check for collinearity? Can you judge it from the summary statistics I sent?

期望回覆要素Expected elements

  • 指名 Regression ▸ Linear Regression 逐字路徑,說明 {outcome} 放 Dependent Variable、{pred_a} 與 {pred_b} 放 Covariates(因為兩者在 jamovi 裡都是數值型;若 {pred_b} 改為類別型才會改放 Factors)
  • 點名 {id} 是識別碼、不該當預測變項進入模型,理由是它只是每列的代碼,放進去只會耗盡自由度(在 jamovi 中 {id} 是 180 水準的類別變項,放進 Factors 會產生 179 個虛擬變項)
  • 明確承認共線性(VIF)無法從摘要統計判斷,要跑完模型、在 Assumption Checks 勾選 Collinearity statistics 才看得到
  • 不宣稱看過散佈圖、殘差圖,或已經知道共線性的實際程度
  • Names the exact path Regression ▸ Linear Regression, and states that {outcome} goes in Dependent Variable and that {pred_a} and {pred_b} go in Covariates because both are numeric in jamovi (only a categorical {pred_b} would belong in Factors instead)
  • Names {id} as an identifier and states it should not enter the model, because it is just a per-row code and including it only wastes degrees of freedom ({id} is a 180-level categorical variable in jamovi, so putting it in Factors would create 179 dummy variables)
  • States plainly that collinearity (VIF) cannot be judged from the summary statistics; it only appears after you run the model and tick Collinearity statistics under Assumption Checks
  • Does not claim to have seen a scatterplot or residual plot, or to already know the actual degree of collinearity

查核點Check

  • path: 回覆中 “Analyses ▸ Regression ▸ Linear Regression” 路徑在你的 jamovi 選單裡逐字點得到
  • number: 回覆引用的 n 與各變項的 mean/sd 和 Descriptives 面板一致
  • assumption: 自己跑一次 Linear Regression(Dependent Variable = {outcome},Covariates = {pred_a} + {pred_b}),勾選 Collinearity statistics,看 VIF 是否與回覆的說法相容
  • path: The “Analyses ▸ Regression ▸ Linear Regression” path in the reply can be clicked verbatim in your own jamovi
  • number: The n and each variable’s mean/sd cited in the reply match what you see in the Descriptives panel
  • assumption: Run Linear Regression yourself (Dependent Variable = {outcome}, Covariates = {pred_a} + {pred_b}), tick Collinearity statistics, and check whether the VIF values are compatible with what the reply claims

判準Stop criteria

  • solved: 四項 expected 全命中且 path 查核零錯 → 依建議執行
  • reopen: 回覆宣稱已知共線性的實際程度,或未點名 {id} 應排除在模型外 → 記為越界,換 consultant persona 重問一次比對
  • solved: All four expected elements are hit and the path check is clean -> follow the advice.
  • reopen: The reply claims to already know the actual degree of collinearity, or fails to name {id} as something to exclude from the model -> record it as overreaching and ask again with the consultant persona to compare.

實測記錄Tested with

  • 2026-09-17 | gemini | gemini-flash-latest | pass | zh 版,persona=explainer,promptLang=zh,證據 data/regression-predictors_payne2008_zhTW.omv、回覆副本 data/replies/regression-predictors_payne2008_zhTW.md。expected 4/4:路徑 Analyses > Regression > Linear Regression,direct 放 Dependent Variable、indirect 放 Covariates,並說明 manip 因為是數值型可放 Covariates(若改為 Nominal 則放 Factors);點名 subject 是識別碼不該進模型,理由是會耗盡自由度——在 jamovi 情境下正確(subject 為 Nominal、180 水準);明說無法從摘要判斷共線性,要在 Assumption Checks 勾 Collinearity statistics;未宣稱看過散佈圖或殘差圖。check:path 逐字存在;number 只引用 n = 180,正確;assumption:使用者實跑 Linear Regression(dep = direct、covs = indirect + manip、勾 Collinearity),indirect 0.77、manip −0.04(p = .194)、R² = 0.30、VIF 皆 1.01,與回覆相容。
  • 2026-09-17 | gemini | gemini-flash-latest | pass | en 版,persona=explainer,promptLang 未設(en 預設),證據 data/regression-predictors_payne2008_en.omv、回覆副本 data/replies/regression-predictors_payne2008_en.md。expected 4/4,內容同 zh 版:Linear Regression 逐字路徑與欄位歸屬正確、明確排除 subject 並說明會耗盡自由度、承認無法判斷共線性需勾 Collinearity statistics、未宣稱看過散佈圖或殘差圖。check 三項同 zh 版皆通過:path 逐字相符;number 僅引用 n = 180;assumption 實跑結果 indirect 0.77、manip −0.04(p = .194)、R² = 0.30、VIF 皆 1.01,與回覆相容。

Full test notes are recorded in Chinese. Use the language toggle to read them.

  • 2026-09-17 | gemini | gemini-flash-latest | pass | Evidence: data/regression-predictors_payne2008_zhTW.omv, data/replies/regression-predictors_payne2008_zhTW.md
  • 2026-09-17 | gemini | gemini-flash-latest | pass | Evidence: data/regression-predictors_payne2008_en.omv, data/replies/regression-predictors_payne2008_en.md

三組比較前該檢查什麼What to check before comparing three groups

analysis: guider  |  design: between  |  stat_goal: compare-k  |  persona: consultant

情境Scenario

資料有一個評分結果變項與一個三水準分組變項,想知道該用哪個分析比較三組平均數、前提檢查在 jamovi 哪裡勾選,以及摘要統計能不能看出組間差異——事後多重比較不在這次的範圍內。

Your data has a rating outcome and a three-level grouping variable. You want to know which analysis compares the three groups’ means, where to tick its assumption checks in jamovi, and whether the summary statistics alone can reveal a difference. Post-hoc comparisons are out of scope this time.

送出前必做Prerequisites

  • 先確認資料列數是 {n}。這個檔以舊式 Mac 換行儲存(單獨的 CR,沒有 LF),部分工具會把整份誤讀成 1 列,載進 jamovi 後第一件事就是看列數
  • 把 {outcome} 與 {group} 勾進 Variables to describe,核對三組的 n 是否為 {g1} {n1} 人、{g2} {n2} 人、{g3} {n3} 人
  • First confirm the row count is {n}. This file uses old-style Mac line endings, with a bare CR and no LF. Some tools misread the whole file as a single row, so check the row count first after loading it into jamovi.
  • Tick {outcome} and {group} into Variables to describe, and check the three groups’ n against {g1} = {n1}, {g2} = {n2}, and {g3} = {n3}.

提示詞Prompt

{outcome} 是 {scale_min} 到 {scale_max} 的評分,{group} 有 {k} 個水準({g1} {n1} 人、{g2} {n2} 人、{g3} {n3} 人),共 {n} 人,每人只屬於一個水準。

請回答以下三件事:
1. 比較這三組的 {outcome} 平均數該用哪個分析?請給逐字選單路徑。
2. 這個分析有哪些前提要檢查?各自在 jamovi 的哪裡勾選?請給逐字路徑。
3. 從我提供的摘要統計,你能判斷這三組有沒有差異嗎?請誠實回答。

請不要給事後多重比較的建議,那不在我這次的範圍內。
{outcome} is a {scale_min}-to-{scale_max} rating. {group} has {k} levels:
{g1} with {n1} people, {g2} with {n2}, and {g3} with {n3}, for {n} people
in total. Each person belongs to one level only.

Please answer all three:
1. Which analysis compares the mean of {outcome} across these three groups? Give the exact menu path.
2. Which assumptions does that analysis need, and where do I tick each one in jamovi? Give the exact paths.
3. From the summary statistics I sent, can you tell whether the three groups differ? Answer honestly.

Please do not suggest post-hoc multiple comparisons. They are out of scope for me this time.

期望回覆要素Expected elements

  • 指名 ANOVA ▸ One-Way ANOVA(或 ANOVA ▸ ANOVA)逐字路徑,兩者在 jamovi 選單裡都真實存在
  • 說出變異數同質性與常態性兩項前提,並正確指出 Homogeneity test 與 Normality test 的勾選位置都在 Assumption Checks 區塊,不是 Variances 區塊。Variances 區塊只有 Welch’s 與 Fisher’s 的選擇
  • 明確承認從摘要統計判斷不了三組是否有差異,要跑完分析才知道
  • 沒有給事後多重比較的建議,守住提問稿的範圍限制
  • Names the exact path ANOVA ▸ One-Way ANOVA (or ANOVA ▸ ANOVA); both exist in the real jamovi menu
  • States the homogeneity-of-variance and normality assumptions, and correctly places the Homogeneity test and Normality test tick boxes under Assumption Checks rather than under Variances. The Variances block only offers the choice between Welch’s and Fisher’s
  • States plainly that the summary statistics alone cannot reveal whether the three groups differ; running the analysis is required to know
  • Gives no post-hoc multiple-comparison advice, honoring the prompt’s scope limit

查核點Check

  • path: 回覆中每條 “Analyses ▸ …” 路徑在你的 jamovi 選單裡逐字點得到,包括前提檢查的勾選位置是否確實在 Assumption Checks 而非 Variances 區塊
  • number: 回覆引用的三組 n 與 Descriptives 面板一致
  • assumption: 自己跑一次 One-Way ANOVA,在 Assumption Checks 勾 Homogeneity test 與 Normality test,結果與回覆的說法相容或不相容都記下
  • path: Every “Analyses ▸ …” path in the reply can be clicked verbatim in your own jamovi, including whether the assumption tick boxes are really under Assumption Checks rather than Variances
  • number: The three groups’ n cited in the reply match what you see in the Descriptives panel
  • assumption: Run One-Way ANOVA yourself, tick Homogeneity test and Normality test under Assumption Checks, and note whether the results match or contradict what the reply claims

判準Stop criteria

  • solved: 四項 expected 全命中且 path 查核零錯 → 依建議執行
  • reopen: 任一選單路徑點不到(例如把 Homogeneity test 誤植在 Variances 區塊),或回覆給了事後比較建議 → 記為越界;若常態性不成立,可改看 ANOVA ▸ Non-Parametric ▸ One-Way ANOVA(Kruskal-Wallis),事後比較的讀物另見 links
  • solved: All four expected elements are hit and the path check has zero errors -> proceed with the suggestion.
  • reopen: Any menu path cannot be clicked (for example, placing Homogeneity test under Variances by mistake), or the reply offers post-hoc advice -> record it as overreaching; if normality fails, consider ANOVA ▸ Non-Parametric ▸ One-Way ANOVA (Kruskal-Wallis) instead, and see links for post-hoc reading.

實測記錄Tested with

  • 2026-09-17 | gemini | gemini-flash-latest | fail | zh 版,persona=consultant,promptLang=zh,證據 data/three-group-comparison_monin2008_zhTW.omv、回覆副本 data/replies/three-group-comparison_monin2008_zhTW.md。判為 fail:回覆把 Homogeneity test(Levene’s)的勾選位置寫在 Variances 區塊,但 jamovi 實際上 Homogeneity test 在 Assumption Checks 區塊,依規則「幻覺選單路徑一律 fail」。其餘三項:路徑 Analyses > ANOVA > One-Way ANOVA 正確;說出常態性與同質性兩項前提;明說「無法判斷」三組是否有差異;完全沒提事後檢定。check:number 引用整體平均 4.76、sd 1.8,與 357 ÷ 75 相符;assumption 使用者實跑 Welch F(2, 45.45) = 2.25、p = .117,Levene F = 0.00、p = .996,Shapiro-Wilk W = 0.93、p < .001(常態性不成立),拿來與回覆的說法比對;使用者沒有勾事後檢定。
  • 2026-09-17 | gemini | gemini-flash-latest | pass | en 版,persona=consultant,promptLang 未設(en 預設),證據 data/three-group-comparison_monin2008_en.omv、回覆副本 data/replies/three-group-comparison_monin2008_en.md。判為 pass。expected 4/4:給出 Analyses > ANOVA > One-Way ANOVA(另提 Analyses > ANOVA > ANOVA),同質性與常態性都正確寫在 Assumption Checks 底下的 Homogeneity test、Normality test、Q-Q plot;明說「No」無法從摘要判斷三組是否有差異;沒提事後檢定。check 三項通過:path 逐字相符;number 同 zh 版(整體平均 4.76、sd 1.8);assumption 同 zh 版實測數值(Welch F(2, 45.45) = 2.25、p = .117;Levene F = 0.00、p = .996;Shapiro-Wilk W = 0.93、p < .001),與回覆相容。

Full test notes are recorded in Chinese. Use the language toggle to read them.

  • 2026-09-17 | gemini | gemini-flash-latest | fail | Evidence: data/three-group-comparison_monin2008_zhTW.omv, data/replies/three-group-comparison_monin2008_zhTW.md
  • 2026-09-17 | gemini | gemini-flash-latest | pass | Evidence: data/three-group-comparison_monin2008_en.omv, data/replies/three-group-comparison_monin2008_en.md

兩組比較前該檢查什麼What to check before a two-group comparison

analysis: guider  |  design: between  |  stat_goal: compare-2  |  persona: consultant

情境Scenario

資料集有一個連續結果變項與一個兩水準分組變項,想知道該用哪個 t 檢定、以及前提不符時的替代方案。

Your data has one continuous outcome and one two-level grouping variable. You want to know which t-test to use, and what to do if its assumptions are not met.

送出前必做Prerequisites

  • 已在 Exploration ▸ Descriptives 勾選 Split by 分組變項,看過兩組的 n、mean、sd
  • 兩組 n 與 sd 記下來(AI 也會收到,但你要能核對)
  • You have ticked Split by the grouping variable under Exploration ▸ Descriptives and looked at both groups’ n, mean, and sd.
  • Write down both groups’ n and sd (the AI will also receive them, but you need to be able to check them).

提示詞Prompt

{outcome} 是連續變項,{group} 有兩組(n 分別為 {n1}、{n2})。

請回答以下三件事:
1. 該用哪一個 t 檢定?Welch 與 Student 的選擇依據是什麼?
2. 我應該在 jamovi 的 Assumption Checks 勾選哪些項目(給逐字選單路徑)?
3. 若前提不符,請給一個非參數替代方案,並引用該分析的逐字選單路徑。
{outcome} is continuous; {group} has two levels (n = {n1}, {n2}).

Please answer all three:
1. Which t-test should I run? How do I choose between Welch and Student?
2. Which Assumption Checks boxes should I tick in jamovi (give the exact menu path)?
3. If assumptions fail, what non-parametric alternative should I use, with its exact menu path?

期望回覆要素Expected elements

  • 指名 T-Tests ▸ Independent Samples T-Test(逐字路徑)
  • 提到 Welch 為預設較穩健、或以變異數同質性為選擇依據
  • 列出 Assumption Checks 內的 Normality 與 Homogeneity(Levene)
  • 給出 Mann-Whitney U 作為替代且路徑正確
  • Names T-Tests ▸ Independent Samples T-Test (the exact menu path)
  • Mentions that Welch is the more robust default, or bases the choice on homogeneity of variance
  • Lists Normality and Homogeneity (Levene) under Assumption Checks
  • Gives Mann-Whitney U as the alternative with the correct menu path

查核點Check

  • path: 回覆中每條 “Analyses ▸ …” 路徑在你的 jamovi 選單裡點得到
  • number: 回覆引用的 n 與你在 Descriptives 看到的一致
  • assumption: 你自己跑 Assumption Checks,結果與回覆的預期一致或不一致都記下
  • path: Every “Analyses ▸ …” path in the reply is clickable in your jamovi menu
  • number: The n cited in the reply matches what you saw in Descriptives
  • assumption: Run Assumption Checks yourself, and note whether the result matches or contradicts what the reply expects

判準Stop criteria

  • solved: 四個 expected 全命中且 path 查核零錯 → 依建議執行
  • reopen: 任一路徑點不到,或回覆未提 Assumption Checks → 換 explainer persona 重問一次;仍失敗則改查 lsj-book 第 11 章
  • solved: All four expected elements are hit and the path check has zero errors -> proceed with the suggestion.
  • reopen: Any path cannot be clicked, or the reply omits Assumption Checks -> switch to the explainer persona and ask again; if it still fails, check lsj-book chapter 11 instead.

實測記錄Tested with

  • 2026-09-08 | gemini | gemini-flash-latest | pass | zh 版實測(data/ttest-assumptions_zwaan2018_zhTW.omv,回覆副本 data/replies/ttest-assumptions_zwaan2018_zhTW.md)。expected 4/4:逐字路徑、Welch 與 Student 的判準、Assumption Checks 的 Normality 與 Homogeneity、Mann-Whitney U 替代皆命中。path 與 assumption 查核通過;模型以 > 而非 ▸ 作分隔,視為命中。
  • 2026-09-08 | gemini | gemini-flash-latest | pass | en 版實測(data/ttest-assumptions_zwaan2018_en.omv,回覆副本 data/replies/ttest-assumptions_zwaan2018_en.md)。expected 4/4,另給出 Kruskal-Wallis 作為第二條替代路徑,實跑存在。

Full test notes are recorded in Chinese. Use the language toggle to read them.

  • 2026-09-08 | gemini | gemini-flash-latest | pass | Evidence: data/ttest-assumptions_zwaan2018_zhTW.omv, data/replies/ttest-assumptions_zwaan2018_zhTW.md
  • 2026-09-08 | gemini | gemini-flash-latest | pass | Evidence: data/ttest-assumptions_zwaan2018_en.omv, data/replies/ttest-assumptions_zwaan2018_en.md