Judgment in Practice Who decides?

What the evidence can tell us.

These sources informed the five claims. Each entry describes what it contributes and where further evidence is needed.

The argument behind the discussion

Jack Shaw · The Irreducible Officer

The essay proposes that purpose, framing, reliance, and accountability must remain visible as AI takes on more work. Its seven patterns organize this discussion. Written for professional military education, it is adapted here for an EduFish education audience.

We adapted the five claims from the essay’s argument for stronger human judgment within capable AI use. Their usefulness as assessment criteria would need to be tested.

For a deeper read: sections II (failure modes), III (finished work), IV (purpose), VI (developmental friction), and VIII (assessment). The companion and faculty workbench offer ways to explore the argument.

Performance with help and learning afterward

Bastani and colleagues · Generative AI without guardrails can harm learning · PNAS, 2025

A randomized study in high-school mathematics found improved practice performance with GPT-4 assistance. Students using the less constrained interface later performed worse without it; teacher-designed tutoring safeguards largely mitigated that harm.

Supports: Distinguishing performance during assistance from later independent performance, and examining how help is designed.

Limit: The researchers measured mathematics performance with particular tools and students. Whether these results extend to other subjects, or to ownership of judgment, remains open.

Which effort develops judgment?

The essay’s developmental-friction argument asks educators to identify which efforts build the capability they want students to develop. Claim 3 applies that distinction to reading and synthesis.

The tutoring study above offers evidence that the design of assistance matters. Educators still have to decide which efforts to preserve for a particular subject and learner, including what support they need to participate.

Framing, authority, and answerability

The Irreducible Officer describes frame capture, invisible delegation, and responsibility laundering. Claims 2 and 4 use these as lenses for noticing choices hidden inside apparently routine assistance.

Eli Talbert · The Triage Trap argues that AI can narrow options before deliberation. This is a military-context argument; its application to schools here is an analogy.

Addy Osmani · Own the Outer Loop develops a practitioner account of evidence, decisions, and answerability in agentic software work. We use its account of who can explain and correct a decision.

Different outputs and diversity of thought

Deng, Brucks, and Toubia · Examining and Addressing Barriers to Diversity in LLM-Generated Ideas · Research preprint, 2026

The authors examine differences in how humans and language models generate diverse ideas and investigate interventions. Targeted prompting also increased diversity in their experiments. The findings suggest that prompt design can affect the range of ideas a group receives.

Limit: The studies examined idea-generation tasks. Applying their findings to a classroom would require examining students’ assumptions and how those assumptions developed.

Accepting useful advice and rejecting poor advice

Raees and Papangelis · From Trust to Appropriate Reliance · Research review, 2026

The review of 22 studies distinguishes feelings of trust, acts of reliance, and whether reliance is warranted. Measures vary. The choice of measure affects what a study can tell us about accepting or challenging advice.

Dell’Acqua and colleagues · Navigating the Jagged Technological Frontier reports task-dependent benefits and failures in a 2023 field experiment with consultants using GPT-4. Applying these findings to newer agents would require fresh testing.

Assessment and authority: further proposals

The tool-invariant assessment framework and authority ladder commentary propose approaches to assessment and delegated authority. Their effectiveness across settings remains to be evaluated.

What this site adds

The five provocations, educational examples, follow-up questions, and facilitator timing are our synthesis. All examples are fictional or illustrative. The 20–30 minute format is a facilitator’s plan. We have yet to test it with an EduFish group or measure what participants learn.

The site is a shared reference for spoken discussion. It collects no responses. External resource links lead to separate sites with their own access and data practices.