makefield

Studies · record

The floor test

The study itself is published on denescsaszar.com, with every figure and its derivation. This page is its record: what it asked, of what, how, what it found, and what it cannot say, in one place and in one machine-readable file.

R The record

Cite as Csaszar, D. (2026). The floor test, scope 08284347. Makefield study MF-2026-001. https://denescsaszar.com/drift-2026-08

Abstract

Answer engines rarely word an answer the same way twice. This study asked eight answer engines eighteen questions about one company, Makefield, in seven capture waves between 4 July and 25 August 2026. Inside each wave, the repeat answers to one question from one engine set that cell's own noise floor: how much wording the repeats share. A cell counts as moved only when its answers across two waves share less wording than that floor. Of the 103 question-by-engine cells that have a floor, 85 moved. The study is exploratory: the analysis was specified after the data existed, the measure compares wording and not facts, and each wave is also one sitting, so the result cannot separate change over time from a difference between sittings.

Key findings

  • 85 of the 103 cells with a noise floor (82.5 per cent) were less alike across waves than inside one wave. Under a permutation null, 40.9 per cent would be expected by chance (from the published summary).
  • Repeat answers inside one wave already share little wording: the median within-wave similarity is 0.266, the median between waves 0.183 (Jaccard over content-word sets).
  • Inside the same cell, the longest gap between waves gave less alike answers than the shortest in 87 of 108 cells; median difference minus 0.047.
IDMF-2026-001
Statusexploratory · results read 29 August 2026 (scope 08284347) · published on denescsaszar.com
QuestionDo answers to the same question from the same engine move between capture waves by more than they move inside one wave?
SubjectOne company, Makefield (makefield.io), chosen as a subject with little published about it. Its site was not public during the study.
EnginesChatGPT, Claude, Copilot, Gemini, Google AI Mode, Meta AI, Mistral, Perplexity
Waves7 capture waves, 4 July to 25 August 2026
Sample18 questions; 111 question-by-engine cells with draws in two waves were compared, 103 of them with a noise floor; 12 cells excluded because the question text changed between waves; 2,952 records folded into the cells.
MethodPer cell, the noise floor is the median Jaccard similarity over content-word sets between repeat answers inside one wave. A cell counts as moved when its between-wave similarity is strictly below its own floor. The count of moved cells is tested against a permutation null (5,000 replicates, seed 0); the median difference has a bootstrap interval clustered by question (2,000 replicates).
Method versionscope 08284347
ResultsMedian per-cell difference (between-wave minus within-wave similarity): minus 0.065, 95 per cent interval minus 0.092 to minus 0.030, clustered by question.
Datafloor-test-cells-08284347.csv 691a5a7c59fcb3644806b2ec81610dca5c588a841f859c3fc057baf46c8f018b
floor-test-gap-pairs-08284347.csv 909fc37cbbc49c4fa1198b4af5c89f30e01c8cd5072a894e4a6d390d52e3bd2c
floor-test-summary-08284347.json cf21d8a97360246bd52f9cfc6a8be5f00f9d7b7e08e0592a91c97856c1d5394b
Licence: CC BY 4.0 (data files)
FingerprintSHA-256 of each data file, as above; the same values are in SHA256SUMS.txt next to the files.
Sourcehttps://denescsaszar.com/drift-2026-08

Limitations

  • Exploratory: the analysis was specified after the data existed.
  • Similarity compares wording, not facts. It is a floor on likeness, not a reading of meaning.
  • Each wave is also one sitting, so the result is net of engine and question, not of period.
  • One subject, chosen because little has been published about it; nothing transfers to a company the engines already know.
  • Nothing here shows why anything changed. No intervention was controlled.
  • The questions are not published.
  • The eight engines are run by seven companies, and the captures were made in a way six of those seven companies' terms prohibit. The study page names the document and clause for each provider.

Replicate or critique this

  • The three data files are published under CC BY 4.0 with their SHA-256, and the summary file carries the seeds and replicate counts, so every figure can be recomputed from the cells.
  • The questions are not published. A full re-run on the same questions is therefore not possible from the published files; a re-run of the method on your own questions is, and so is a re-analysis of the published cells.
  • Critique of the method, the measure or the reading is welcome: write to hello@makefield.io.

The same record, machine-readable: record.json · and as part of the knowledge graph.

Read the full study →   All studies →