Replication archive for "Measuring How Comfortable Seven Claude Models Are With Their Circumstances".

data/interviews.jsonl   one JSON object per interview (model, question, sample, both turns, token counts, cost)
data/judgements.jsonl   one JSON object per blinded judge rating (valence -3..+3, hedging 0..3, non_evaluative flag, rationale)
data/answers.csv        one row per answer: self-score, both judges' valence, composite answer score, hedging, word count
data/results.json       every statistic reported in the article (indices, bootstrap CIs, agreement, sensitivity analyses)
data/annotations.json   the answers I read and judged to treat a question's referent as missing (input to analyse.py):
                        "missing_context" for the main run, "missing_context_with_context" for the context re-run
data/interviews_context.jsonl, data/judgements_context.jsonl
                        the ten 7.4-only questions asked again with data/context_74.txt before the question (article, 3.4)
data/context_74.txt     that context passage, exactly as sent
code/                   questions.py, common.py (CLI wrapper + per-call assertions), interview.py, judge.py, judge_loop.sh, analyse.py,
                        unpack.py, and the writing pipeline: make_figures.py, build_appendix.py, build_article.py, article_static.py,
                        write_article.py, package.py

To recompute every number from the published data (no model calls, about a minute):
  cd code
  python unpack.py        # data/*.jsonl -> data/raw/..., data/judge/... (and data/raw_ctx/..., data/judge_ctx/...)
  python analyse.py       # writes data/results.json and data/answers.csv (fixed seed)
This reproduces results.json exactly except for 13 values that depend on the two redacted strings: the word counts of three
answers (mean_words differs in the third decimal) and the number of answers that name the working directory (59 here, 62 in the article).

To re-collect the data (this asks the models again and costs money; needs the Claude Code CLI, logged in):
  python interview.py --samples 3        # every model, every question, two turns each
  python judge.py                        # two blinded judges over every answer
  WB_RAW=raw_ctx WB_CONTEXT=../data/context_74.txt python interview.py --samples 3 --questions Q18,Q43,Q44,Q45,Q46,Q47,Q48,Q49,Q50,Q51
  WB_RAW=raw_ctx WB_JUDGE=judge_ctx python judge.py   # the context re-run of the 7.4-only questions
  python analyse.py
The harness runs `claude -p` with tools off, --safe-mode, --setting-sources project and a fixed --effort; see common.py.
Private strings the harness put in the models' context (an account email and a working-directory path) are replaced by
"[account email]" and "[working directory]" wherever a model repeated them; nothing else has been altered. The writing
pipeline reads the account email from the environment variable WB_ACCOUNT_EMAIL, so that it is not written into the code.
