Zak FentonWorking notes

· Note

What checking my research proposal found

What the checking found, and why this site shows its working

I had a proposal for a side project checked. The audit logged 92 entries: 26 labelled confirmed, 27 corrected, 36 unsupported and three unverifiable. Those labels were messier than they looked. Some “confirmed” entries still needed changes.

None of it was made up on purpose. All of it would have been found by anyone who opened the right paper. I should have opened those papers before using them.

What the check found

I’d used Risko and Gilbert’s review to support a figure about AI-assisted work. The 40 per cent passage describes how often people chose to write down two items in a memory task. It says nothing about AI-assisted work.1

I’d listed a manuscript as a CHI 2016 conference paper. The document I found was still in a submission template. It reported no significant difference in exercise behaviour between the groups.2

I’d called a finding significant when the study reported no statistical test at all, had no bias task, and I’d got the journal’s name wrong.3

I’d written that a study would have 80 per cent power. Recomputed, it was 41.6 per cent.

One inequality was written backwards, so the hypothesis asserted the opposite of my own argument.

The audit also pointed me to a review whose reference list included a paper claiming AI makes people lazy. That paper was retracted on 3 February 2026.4

What broke

I still wanted to pursue the idea. The proposal needed changes to its sources and research design. Most of the damage was sourcing: claims lifted from summaries of papers instead of the papers, and numbers carried over instead of recomputed.

Honestly, that’s the ordinary way errors get into evidence. Fluent summaries make it faster, because a confident paragraph reads exactly like a checked one.

The corrected proposal was better than the original. Annoying, but true.

Why this site works the way it does

The evidence reviews here don’t ask to be trusted. Every assessed claim shows the exact passage from its source, what that source can’t tell you, and the date it was checked. If a source couldn’t be reached, the entry says what was tried and why it failed. A claim that can’t show its working stays in review, with no verdict. When I get something wrong, the correction is dated and listed.

It’s slower. Frankly, it’s the only version I’d put my name on.

The receipt

The audit is dated 28 July 2026 and records five verification tracks and a red-team pass. It reports primary-source reads where access allowed and recomputation of its headline numbers. The full pack stays private, because it’s a proposal for a side project and the proposal isn’t the point. The audit’s recorded categories:

  • Confirmed: 26
  • Corrected, where a defensible version of the claim survived: 27
  • Unsupported by the source cited or by any source the audit found: 36
  • Could not be verified: 3

These are the audit’s labels, not a count of claims that passed without changes. Its confirmed group includes two entries that needed qualification and another that was a suggested addition to the proposal.

The examples above are audit rows 12 (the 40 per cent figure), 65 and 68 (the conference paper), 86 (the significance claim), 48 (the power calculation), 51 (the inverted inequality) and 7 (the retracted paper). The power figure assumes 20 participants, a population correlation of 0.40, alpha of 0.05 two-sided, and the Fisher z approximation: 0.4157, or 41.6 per cent. It is a planning approximation, not a measured result. The review the audit pointed to for the retracted paper is Zhai, C., Wibowo, S. and Li, L. (2024). The effects of over-reliance on AI dialogue systems on students’ cognitive abilities: a systematic review. Smart Learning Environments, 11, 28. doi:10.1186/s40561-024-00316-7.

Footnotes

  1. Risko, E. F. and Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9), 676-688. doi:10.1016/j.tics.2016.07.002.

  2. Ramsay, D. B., Jin, J., Maes, P. and Picard, R. Virtual pets and virtual selves as exercise motivation tools. Unpublished manuscript, MIT Media Lab and MIT Sloan School, undated, author-hosted PDF. Its abstract gives 20 participants and its methods 21, so no participant count is stated here.

  3. Ben-Zion, Z. et al. (2025). Assessing and alleviating state anxiety in large language models. npj Digital Medicine, 8, 132. doi:10.1038/s41746-025-01512-6.

  4. Ahmad, S. F., Han, H., Alam, M. M., Rehmat, M. K., Irshad, M., Arraño-Muñoz, M. and Ariza-Montes, A. (2023). Impact of artificial intelligence on human loss in decision making, laziness and safety in education. Humanities and Social Sciences Communications, 10, 311. doi:10.1057/s41599-023-01787-8. Retraction note: Humanities and Social Sciences Communications, 13, 150, 3 February 2026. doi:10.1057/s41599-026-06602-8. Dates and details from the publisher’s Crossref record, 15 September 2026; the article pages were behind a login when checked.