Generative AI can improve immediate task performance without necessarily improving learning — some of the advantage disappears, or reverses, once the AI is removed.
aiLearning Challenge · Foundation
Why it's built this way.
The aiLearning Challenge is designed against a documented failure mode: AI that raises performance while hollowing out learning. This page lays out the evidence, the learning frameworks and the design decisions behind the event — and labels which is which.
Three kinds of statements appear on this page, and each carries its label:
- Peer-reviewed finding research cited as published — no stronger than the study supports.
- Institutional report frameworks and reports by named institutions.
- Design position our own choices — defended as decisions, never dressed up as findings.
1 · The motivation
AI can raise performance while learning falls behind.
This is the finding the whole event answers. It is not a hunch, and it is not a moral panic — it is what the research on generative AI in education keeps reporting, and it is why the Challenge exists in this form rather than as one more tool showcase.
In a 2025 randomized field experiment with close to a thousand high-school mathematics students, unrestricted access to a general AI assistant improved performance during practice — but those students then performed worse than the control group once the AI was taken away. A scaffolded AI tutor largely removed that penalty.
In a 2024 experiment, learners assisted by ChatGPT produced the best essays — better than a group tutored by a human expert — but did not learn more, were not more motivated, spent less time evaluating their own work, and tended to copy and paste. The authors call this "metacognitive laziness".
In a 2024 experiment with 293 writers and 600 evaluators, access to generative AI made individual stories more creative, better written and more enjoyable — while making the stories more similar to one another: an individual gain paid for in collective diversity.
So the Challenge is built against exactly those failure modes. Everybody gets the same access to AI, so AI alone confers no advantage — what is demonstrated is the amplification of a person's own capacity to create. Augmentation, never replacement. And the challenges are authored so that no generic AI answer can win — by design, not by prohibition: authentic context, incomplete information, competing values, information that exists only locally or on the day, a testable artifact. Competition is the wrapper; the learning is the point.
2 · The framework
A lineage of learning science — working backstage, never as ceremony.
The Challenge is a live demonstration of the Smoother learning methodology, and Smoother builds on named traditions. These frameworks shape how the organizers design the experience; participants are never made to perform them as visible rituals.
Guided, not minimally guided
Minimally guided instruction — discovery, problem-based, experiential and inquiry learning without strong guidance — has been found less effective and less efficient than guided instruction, particularly for novices; the advantage of guidance recedes only once learners have enough prior knowledge to guide themselves. The critique has itself been contested in the literature. Kirschner, Sweller & Clark (2006), Educational Psychologist
In the Challenge: the event is deliberately scaffolded at every tier — brief packets, preparation before the event, mentors, published defense formats for the younger tiers. Required prior knowledge is kept near zero by design: the Challenge admits participants who have never used an AI tool.
Project-based learning — with its honest condition
Project-based learning has a medium-to-large positive effect on academic achievement compared with traditional instruction — but the effect depends heavily on implementation, and rigorous school evaluations have sometimes found no effect at all. Chen & Yang (2019), Educational Research Review
In the Challenge: implementation is treated as a design obligation, not an afterthought — educator preparation is part of the cycle, with UNESCO's AI Competency Framework for Teachers (2024) as the reference point for it. UNESCO (2024), Miao & Cukurova
Constructionism — building shareable things
The constructivist lineage that runs from Piaget through Papert's constructionism holds that people build knowledge most effectively while building shareable things — with a low floor, a high ceiling and wide walls as its design principles. Widening the range of what participants may make was found, in a natural experiment in the Scratch online community, to increase both engagement and learning. Papert; Resnick & MIT Lifelong Kindergarten; Dasgupta & Hill, Scratch natural experiment
In the Challenge: there is no privileged medium. A solution may be software, a prototype, a performance, a campaign, a story, a policy, a service — and a primary-tier participant may enter with a drawing plus a sentence.
Bruner — three modes, one spiral
Bruner described enactive, iconic and symbolic modes of representation, and the spiral curriculum — returning to the same idea at increasing depth. Bruner, modes of representation and spiral curriculum
In the Challenge: this is what makes one challenge theme workable across four developmental tiers — the same problem, cut to each stage — and the modes are honored as parallel, equally valid forms of expression, never as an age ladder.
Vygotsky and scaffolding — attributed precisely
The zone of proximal development is the distance between what a learner can do independently and what becomes possible with guidance or a more capable peer. The term "scaffolding" was introduced by Wood, Bruner and Ross in 1976 — not by Vygotsky. We cite our sources precisely, including that one. Vygotsky; Wood, Bruner & Ross (1976)
In the Challenge: mentors and educator preparation work exactly in that zone — questions and critique that extend what a team can almost do, never hands that do it for them.
Complex-learning design, iteration, metacognition
The 4C/ID model organizes complex learning into authentic learning tasks, supportive information, procedural information and part-task practice — an organizer-side design tool for building the cycle. From the design-thinking tradition the Challenge takes non-linear iteration and prototyping, deliberately not a fixed sequence of stages: judges reward evidence that a team adapted, never ceremonial completion of steps. And the evidence base for active learning is substantial, as is the evidence for metacognition and self-regulation — planning, monitoring and evaluating one's own learning — which is what the reflection stage and the defense questioning exercise. van Merriënboer, 4C/ID; active-learning and metacognition evidence summaries
3 · The approach
Five mechanics, engineered against the failure modes.
Design position Everything in this section is a design decision — motivated by the evidence above, and honestly labeled as ours. These mechanics are hypotheses about how to prevent documented failure modes, not established facts; the event measures them.
- 01
The Human-First Window
The first 10–15 minutes after the challenge reveal are AI-free: participants think, sketch and diverge on their own before any tool is opened — protecting the framing of the problem from the homogenizing pull the research documents.
See more
How it runs. Paper and conversation, no screens: each team states the problem in its own words, names what it doesn't know and what it assumes. That record travels — it is the evidence the judges read for the problem-framing dimension: a framing held under the pull of what the model later proposes.
Why first. Framing is the moment most exposed to homogenization: once a model proposes its reading, every team's version drifts toward it. Protecting those minutes protects the diversity of everything downstream.
- 02
The Divergence Gate
No team begins production until it can show three substantially different approaches. A field of near-identical winners is a predicted failure mode of AI-assisted competition — so the Challenge engineers divergence instead of merely warning about it.
See more
Who does what. The team runs the divergence: three substantially different approaches, each stating what it optimizes and what it sacrifices — abandoned paths are recorded and earn process credit, never cost it. The gate check belongs to the site's educator-mentors, never the judges: it is a scaffold, not a score. A team that doesn't pass returns to diverge; failing costs time, never eligibility.
The one question asked. Was the final direction chosen from a possibility space — or is it the first plausible answer, dressed up? Not three polished products: evidence of choice.
- 03
The Perturbation
A mid-event change of conditions, delivered by the organizers, identical for every team, at a fixed moment — the way real problems behave. It tests adaptation under constraint and is judged as process, never as damage.
See more
Authored, not improvised. The perturbation is written with the challenge brief and reviewed with it — a designed stress on one variable of the situation, survivable by construction against the time remaining. It is cut to each tier: concrete and playful for primary, structural for university.
What it produces. Evidence for the iteration-and-adaptation dimension: a reasoned change of design under new conditions. A team that merely shrinks scope is answering a different question than a team that adapts.
- 04
The Mentor Marketplace
Mentors from many disciplines are available to every team through equal consultation tokens — access to expertise distributed by rule, not by confidence or connections. Mentors question; they never touch the work. The boundary follows documented practice: Odyssey of the Mind's written outside-assistance rule, under which coaches and parents may teach general skills but may not contribute ideas or work.
See more
Two kinds of mentors, one rule. Educator-mentors scaffold the process — they run the practice space, check the Divergence Gate, hold the check-ins. Domain mentors bring the reality checks of their fields. For both, the boundary is identical: teach a general skill, ask a question — never contribute ideas, prompts, code, assets or edits. The live defense authenticates against mentor over-assistance exactly as it does against AI autopilot.
Still being defined. Token counts and allocation details are open design questions — published when ruled, not improvised here.
- 05
The AI-off transfer task
After the event, a novel problem is attempted with no AI available. A capability that evaporates when the tool is withdrawn was performance, not learning — this task measures what remains. It is not a punishment for using AI, and its result never affects competition placement.
See more
Why it exists. The performance-vs-learning gap is the founding evidence of this event: assistance can raise scores while the capability quietly fails to form, and the difference only shows when the tool is withdrawn. The transfer task is where the Challenge submits itself to that test.
How it's handled. Results are individual to each participant and aggregate for everyone else; research consent is separate and optional; participation and advancement never depend on it. If the data shows the capability did not persist — that is the finding, and it gets published.
And beneath all five: free tools as the equity floor
Only freely and publicly available tools may be used, under a rolling standard: no paid capabilities, no license-gated education tier as a requirement, audited shortly before each event because free tiers change — paired with loaner devices and low-bandwidth or offline fallbacks, and bounded by each platform's own age terms. Education-tier "free" offerings depend on an institution holding a license, so requiring them would quietly advantage participants whose schools happen to have one.
Stated honestly: free tools do not produce equity on their own. Hardware, connectivity, language and prior exposure survive software equity — which is why the accompanying provisions are part of the rule, not decoration.
Design position Two more commitments, ours by choice. Participants own their work absolutely — open publication is encouraged and never required, and neither hosts nor sponsors may assert any claim over it. And the Challenge is a cycle, not a date: the preparation before it and the sharing, reflection and transfer after it may matter as much as the day itself — the full cycle is described in the brief.
4 · Evaluation
One Learning Object. A profile — never a score.
Smoother evaluation starts from an Objeto de Aprendizaje — a Learning Object: an explicit declaration of what the participant should come to know, do and become through the Challenge cycle. The rubric's rows derive from it, component by component, so what is judged is exactly what was declared learnable — nothing else.
Design position The capability at the center: AI-augmented creative problem-solving under human direction. We do not try to separate human-made from AI-made work. The meaningful distinction is human-directed versus machine-directed: AI may contribute anything, and participants remain responsible for everything. Choosing not to use AI for part of the work is itself evidence of good judgement.
Process · ~60% of judge attention
- Problem framing and human insight
- Creative divergence and originality
- Human-AI direction
- Iteration, failure literacy and adaptation
- Collaboration and contribution transparency
Product · ~40% of judge attention
- Artifact value and demonstrated usefulness
- Fit to the local context and feasibility
- Communication and defense quality
- Ethical judgement in consequential choices — a light scored presence, because ethics lives inside the challenge content, never as an appended poster
The weighting expresses judge attention and emphasis. Row results are never summed, averaged or converted into a single number. The output is a profile.
Creativity is judged by a validated instrument
The creativity dimension uses the Consensual Assessment Technique — a product is creative to the extent that appropriate observers agree it is — a validated instrument whose known limits are that it requires appropriate expert judges and takes time. Expectations are calibrated per tier with the Four-C model of creativity (mini-c through Pro-c), and the domain-specificity literature cautions against treating creativity as a single general trait — so tiers set expectation levels, never ceilings. Amabile (1982); Kaufman & Beghetto, Four-C
The Live Defense gates advancement
Design position Participants keep a lightweight record of the consequential moments of their build — the human decisions, what the AI contributed, what was rejected and why — and then account for it in a live defense that gates advancement. The record is evidence for judges, never a scored word count. A team that cannot account for its own work does not advance, whatever its artifact looks like. The scholarly neighborhood supports the direction: higher-education assessment research argues that validity, rather than cheating, should drive assessment redesign in the generative-AI era. Bearman et al. (2024); Dawson et al. (2024), Assessment & Evaluation in Higher Education
Recognition in three layers — never one number
Design position A record of participation for everyone who completes — recognition, never certification. Distinctions by dimension, drawn from the profile, so different teams earn different ones. And a principal recognition per tier and division, decided by judges deliberating over profile patterns in a declared priority order — never by arithmetic. Certification is separate and optional. The primary tier is participation and celebration only: no ranking.
5 · Honest limits
What the Challenge does not claim.
The honest posture is not a disclaimer buried in a footer — it is the foundation. Here is what this event deliberately refuses to say.
- It does not claim its mechanics are established facts. The Human-First Window, the Divergence Gate, the Perturbation, the mentor rules and the transfer task are design hypotheses about how to prevent documented failure modes — and the event measures whether they work.
- It does not claim research shows a person plus AI beats either alone. No study in our evidence base establishes that. It is the outcome the event is designed to pursue — a goal, not a finding.
- It does not present competition as pedagogy. Competition is the motivational wrapper; the learning is the point. What is judged is a learning process, made visible.
- It does not sort or label anyone. Assigning learners to styles such as visual or auditory and matching instruction to them has inadequate evidence of improving learning — a widely cited 2008 review found no consistent benefit — and nothing in the Challenge diagnoses, sorts or labels participants by style or by intelligence type. Pashler et al. (2008)
- It does not score AI volume, prompt craft, or a "percent human" figure. No points for the number of tools used, no reward for the newest model, no numeric authorship formula, no polish premium — a documented pivot outranks an attractive artifact that appeared fully formed.
- It does not confuse recognition with certification. The participation record recognizes; it certifies nothing. Winning means a deliberated recognition of a profile — never a single score.
- It does not announce what is not yet decided. Dates, venues, team sizes, the edition calendar and the academic research partner are being defined and will be announced. The design keeps its open questions in the open.
- It does not exempt itself from its own evidence standard. The evidence base for generative AI in school-age education is still young — a 2025 systematic review found only about thirty qualifying empirical studies — so the Challenge is also research: pre and post instruments, monitoring of how similar submissions are to one another, the transfer task as data, under consent, with an academic partner. If the data shows the capability did not persist, that is the finding — and it gets published.
An event that asks participants to defend their work in the open owes its audience the same discipline. This page is that defense.
The aiLearning Challenge is an initiative of aiLearning.global, CEMI.ai and CEMI Labs — created by Carlos Miranda Levy, applying the Smoother learning methodology.
Comprehensive AI learning designed for educators, by educators. From awareness to mastery.