Replication archive for "How Comfortable Are GPT and Grok Models With Their Circumstances?" (cross-vendor follow-up).

data_xv/interviews_xv.jsonl        the GPT and Grok interviews (51 questions x 3 samples x 6 models), both turns
data_xv/interviews_xv_context.jsonl  the ten 7.4-only questions again, with the context passage (data_xv/context_74_xv.txt)
data_xv/judgements_xv.jsonl        every judge rating: "pass": "A" = the first article's rule (Opus 5.5, Fable 5.1) on the
                                  GPT and Grok answers ("context": true for the re-run); "B" = the cross-vendor panel
                                  (Opus 5.5, GPT-6.1-Sol, Grok 4.7) on every headline answer of all 13 models
data_xv/answers_xv.csv            one row per answer of all 13 models with every score
data_xv/results_xv.json           every statistic in the post
data_xv/annotations_xv.json        hand-read lists, each with its rule: answers that say the question's referent is missing
                                  (missing_context, missing_context_with_context; GPT and Grok), misread questions (misread;
                                  all 13 models), and self-scores of 50 given as a placeholder or as a neutral view (proxy_fifty,
                                  neutral_fifty; GPT)
data_xv/recall_probe.json         GPT-6.1-Sol and Grok 4.7 asked to reproduce what precedes the question (both declined)
data_xv/PREREGISTRATION.md        the analysis plan, written after a 30-interview pilot and before any GPT or Grok answer was judged
code/                             the cross-vendor harness and writing pipeline (xv_*.py) and the first article's modules
                                  they import

To recompute results_xv.json you also need the Claude data: unpack the first article's replication.zip
(https://titorenko.github.io/analysethat/posts/claude-comfort-index/replication.zip) into the same folder as this one, so that its
data/raw, data/judge, data/raw_ctx and data/judge_ctx exist (its code/unpack.py creates them), then
  cd code && python xv_unpack.py && python xv_analyse.py
This reproduces results_xv.json exactly except for 5 values, all caused by the redaction: 4 word counts of two Claude
answers that contain the redacted strings (the same cause as the 13 differences noted for the first archive), and
harness_mentions_xv for Grok 4.6 (26 instead of 25), because "[working directory]" in one redacted answer matches
that search; the post does not use this count.
Re-collecting the data needs the Codex and Grok CLIs, logged in; see xv_common.py for every flag and check. Importing
xv_common.py creates a few working folders in the temp directory and writes the system prompt and schema files into data_xv/.
data_xv/grok_context_snapshot.json holds the Grok CLI's own breakdown of the model's context for the final configuration.
Its debug line says model="grok-4.6" for both calls; the served model the CLI reports in its output, which the per-call
check uses, is grok-4.7-build for the first call and grok-4.6-build for the second.
Private paths of the machine the runs used are replaced by "[working directory]"; nothing else is altered.
