How the coding essay was actually made
Before I published “Should We Still Teach Coding?”, I did something I would recommend to anyone writing under their own name in this era. I wrote the essay first — my argument, my experience, my point of view — and only then did I ask three separate AI deep-research tools, ChatGPT, Claude, and Gemini, to do one job: verify my statements and find real, citable sources for the things I claim. Not to write it. Not to have opinions. To check me, and to hand me references I could stand behind.
They did that well, and I want to be fair about it — this is augmentation working exactly as it should. They largely confirmed what I had written, and the sources they surfaced are the ones now listed at the end of that essay. Between them they suggested only a handful of corrections, and every one of them made the piece better and more honest:
- They caught that “digital fluency” — the term I was celebrating — was popularized by Mitchel Resnick, the mind behind Scratch. So I credited him. It strengthened the piece; the ending now ties back to its beginning.
- They pushed me to tone down one claim about skill transfer so it reads as lived experience rather than settled science — because the research shows near transfer is strong but far transfer is only modest. So my sentence became “in my experience,” with the honest caveat attached.
- They corrected a framing about Code.org’s rebrand to CodeAI: it did not abandon computer science, it folded it inside “digital fluency.” I fixed it.
- And they let me anchor the calculator analogy in real evidence — with the crucial nuance that the benefit appeared when the tool was integrated into teaching, while dumping it on children too early could set them back.
Four small, true improvements on a piece I had already written carefully. That is the good case. That is what these tools are for.
Then Gemini offered to make an infographic
When Gemini finished its research, it offered something extra: it could turn the findings into an infographic. I said yes, out of curiosity. What came back was genuinely impressive to look at — a clean four-part visual, charts, a confident headline, professional polish. For about a minute I thought about publishing it.
But it did not feel like mine. It was cold. It was statsy — a wall of percentages and risk charts, an alarmist tone about a “vibe coding hangover” and “cognitive debt,” framed to make you anxious. My essay is not anxious. It argues that the foundation endures, that AI augments rather than replaces, that the task is to change education well. The infographic had drifted from my voice into someone else’s.
So I did what we always do. I handed it to our own AI — Claude Code, aligned to my digital twin and persona and to our published guidelines — and asked it to check the infographic against how we actually work. It came back with a page of surprises.
It was not just the tone. Under the polish were real problems: a “developer adoption” statistic with no verifiable source; a claim that a coding-club network operates in “180 countries” when the verified figure is about a hundred; and — the one that would have embarrassed me — invented chart data presented as if it came from named studies, alongside preliminary, non-peer-reviewed findings displayed as settled fact.
None of this means Gemini is a bad tool, or that the research was bad — the underlying dossier was careful and mostly excellent. It means the output drifted: from verified to embellished, from my frame to a generic one, from “sourced” to “sourced-looking.” And it drifted precisely where drift is most dangerous — in the confident, decorative surface that readers trust most.
Drift is normal. Plan for it.
Here is the thing the industry has learned and that individuals using AI often have not: even when you give a model explicit guidelines, structures, and formats, it will still deviate from them. We call it drift. It shows up most on complex requests, and — counter-intuitively — on requests that are only slightly different from what you usually ask, because the model pattern-matches to the familiar and quietly drops the part that was new.
So in our systems, every workflow — no matter how simple or how complex — ends with a compliance checklist. The AI takes its own output and verifies it, item by item, against our guidelines and our approved structures. Whatever fails is regenerated; if enough fails, the whole thing is generated again. It is a self-correcting loop, and I will tell you what surprised us most about it: it catches drift far more often than you would expect. Missing elements, a section that quietly reverted to a default format, a figure that crept in without a source, a tone that slid. The checklist finds it, and the loop fixes it. That discipline is a large part of why we can move fast without publishing things we regret.
If you take one practical thing from this piece, let it be this: if you generate content with AI and you do not have a checking mechanism, build one — or watch every output like a hawk. A written checklist the AI runs against itself, and a human who signs off, is cheap insurance against the one bad artifact that undoes a lot of good work.
What I did instead
I did not publish the infographic. We took the good bones — the real structure, the genuinely verified numbers — and I asked us to rebuild it properly: on our palette, in our voice, self-contained, dark-mode aware, with every figure sourced and every preliminary study labeled as preliminary. You can see that reworked version here — the one that passed the checklist.
And then I wrote this. Because the most useful thing I can share is not the infographic; it is the decision not to ship the first one. That decision — to check, to catch the drift, to choose alignment over polish — is not a tax on working with AI. It is working with AI. The tools are extraordinary. The judgment about whether what they produced is true, and whether it sounds like you, stays with you. It always did.
For the record: this article went through the same checklist — voice, claims, and sources checked before it was published. — Carlos Miranda Levy
Four perspectives
The mechanism Carlos describes has a name in the literature too: verification is not a step you add at the end, it is the thing that makes generative output trustworthy at all. The failure mode in the infographic — invented values displayed with the authority of measured data — is the single most common way AI content misleads, because the chart form itself signals rigor. A compliance checklist that separates 'sourced' from 'sourced-looking,' and that flags preliminary evidence as preliminary, is exactly the right instrument. I would add one item to it: every number must resolve to a source a reader could open.
What worries me is the person without this apparatus. A large organization can build compliance loops; a teacher, a small nonprofit, a student cannot easily. The polished-but-wrong artifact is a genuine equity problem — it is most dangerous for those least equipped to audit it, and it arrives looking more authoritative than the careful, plainer work it competes with. The honest move here is not just to build the checklist for ourselves but to make the checklist itself shareable, so that the discipline is not a privilege of the well-resourced.
Practical version: never let AI output go straight to publish. Put one gate between generation and the world — even a five-line checklist taped to your monitor beats nothing. Does every stat have a link? Does it sound like us? Did it keep the part of the request that was new? Regenerate what fails. The teams that ship fast and clean are not the ones that trust the model; they are the ones that built the cheap, boring gate and never skip it.
I have said for years that the value of AI is not what it does instead of you, but what you can do with it that you never could before. This is the shadow side of the same truth: the more capable the tool, the more convincing its mistakes, and the more the judgment has to remain yours. We caught this one because we treat drift as normal and check for it by default — not because we are clever, but because we are disciplined about it. The infographic was competent. It was not true, and it was not mine. Both of those had to be fixed before it could carry my name. That is the whole job.