Understand where learning breaks. With evidence, not guesses.

PROJECT: CLARION
YEAR: 2026
TYPE: AI-ASSISTED DIAGNOSIC WORKSPACE · EDTECH

00.1 OPENING AND TLDR
Most AI tools fail teachers not by being wrong, but by being too loud.
I designed Clarion around a different principle: an AI earns trust by knowing its limits.
This is the full account of how I built that system, every research finding, every structural decision, and every feature I deliberately left out.
TL;DR
Teachers spend 3–6 hours weekly diagnosing student learning manually.
I designed an AI workspace that earns trust by surfacing only what it can prove, and staying silent when it can't.
Result:
-Weekly diagnostic orientation under 2 minutes.
-Zero auto-approved insights.
-Every decision owned by the teacher, enforced structurally.
CHAPTER 02 · RESEARCH
I wanted to understand what would make this product genuinely useful in a real classroom, not just theoretically valuable.
Since direct access to teachers was limited, I conducted AI-assisted desk research using Perplexity Deep Research, synthesizing insights from educator interviews, teaching forums, academic studies, and classroom management resources.
The research focused on Grade 6–8 teachers across Math, Science, and Language Arts, uncovering recurring challenges around classroom attention, student engagement, workload management, and assessment efficiency.

02.1 RESEARCH FINDING TABLE
METRIC
Weekly diagnostic time
3–6 hours, evenings and weekends
Primary tool failure
No tool detects class-level learning gaps — only submissions and grades
AI tool experience
70% tried AI tools; most stopped due to lack of visible reasoning
Evidence requirement
70% need to see the evidence before trusting any conclusion
Paper prevalence
60% primarily paper-based; full upload requirement is a dealbreaker
Student AI usage awareness
All aware; most uncertain how to respond diagnostically
02.2 THE 4 DIMENSIONS I IDENTIFIED
01
TIME-COLLAPSED PATTERN RECOGNITION
Entirely manual. Entirely memory-dependent. Teachers hold patterns in their heads, and the conclusions may be wrong by Monday.
02
GRADES ARE THE WRONG SIGNAL
One score. Five possible reasons behind it. A grade can't tell you which one, so teachers build the workaround manually.
03
PAPER CREATES AN INVISIBLE WALL
60% of classrooms run on paper. Any system that demands full digitization first fails before it starts.
04
AI WITHOUT EVIDENCE DESTROYS TRUST
Black box in, trust out. If teachers can't see the reasoning, they can't apply their own judgment to it.
"I don't need another grading tool. I need help understanding where my instruction failed so I can fix it."
"I won't trust a system that just tells me what's wrong without showing me why. I need to see the actual student work."
02.3 THE RESEARCH-TO-DESIGN STRATEGY MAP
RESEARCH EVIDENCE
CORE INSIGHT
DESIGN DECISION
3–6 hrs/week manual diagnosis
Diagnosis is the real bottleneck, not grading
Weekly cadence; max 5 surfaced insights
70% need visible evidence
Evidence is a precondition, not a preference
Evidence-first architecture; draft insights only
60% paper-based classrooms
Upload burden kills adoption
Sampling model; no completion pressure
All tools track logistics, not understanding
No tool addresses class-level concept gaps
Diagnostic workspace, not LMS replacement
Black-box AI distrust
Transparency is not a feature — it is trust
Approve/edit/reject; no auto-approval ever
03.1 WHAT I GOT WRONG FIRST
What didn't work
My first design direction included a real-time alert system. When Clarion detected a pattern with high confidence, it would notify the teacher immediately push notification, in-app badge, the works.
Research killed it.
My first direction included real-time alerts. Research killed it.
Teachers called it anxiety-inducing. Impossible to act on mid-lesson.
One teacher put it plainly: "I don't need my phone telling me something is wrong while I'm standing in front of 35 students."
I removed the entire real-time layer. Not deprioritised. Removed.
04.1 PRODUCT VISION
What Clarion Is and What It Is Not
Before I drew a single wireframe, I spent time writing down what Clarion was not. In educational AI, wrong product positioning is not just a marketing error, it is an ethical failure.

04.2 IS / IS NOT TABLE
CLARION IS
A weekly diagnostic workspace
A class-level pattern detector
An evidence surfacing system
A human-in-the-loop aid
A sampling-based signal system
A teacher authority preserver

04.3 5 NON-NEGOTIABLE PRINCIPLES
05.1 THE AUTHORITY GOVERNANCE MODEL
Where AI can appear, and where it is forbidden
AI systems fail in education not because of poor models. They fail because of authority creep, the gradual drift from assistant to advisor to authority. I designed explicitly against this failure mode.
I defined every surface where AI is permitted. I locked every surface where it is forbidden. This is system law, not UX guidance.
AI PERMITTED IN
AI FORBIDDEN FROM
Draft insight text blocks — conditional language only
Evidence annotations — descriptive only, no interpretation
Summaries or conclusions
Withdrawal messages — insufficient evidence trigger
Navigation labels
Silence states — absence of output with neutral reason
Calls to action
Boundary disclosures — what the system cannot do
Empty states
Error states — technical limitation descriptions only
Comparative or temporal frames
Teacher notes or reflections
05.2 THE FOUR SYSTEM LAWS
Upload ≠ Analysis
The system doesn't start watching you the moment you give it data
Insight ≠ Action
Seeing something doesn't mean the system acts on it
Analysis ≠ Insight
Not everything the system detects gets shown to you
Teacher decides every transition
Every step requires a conscious human choiceystem doesn't start watching you the moment you give it data
Any appearance of AI output outside the permitted surfaces is an authority breach, regardless of tone, usefulness, or intent.
06.1 THE FOUR-LAYER IA
Structure as authority
Most product teams move from research to wireframes. I introduced an intermediate layer: a governed information architecture that encoded authority restraint structurally, not visually.
The reasoning: if the hierarchy is wrong, language restraint will fail. Copy can be edited. Structure is harder to break.
Every screen in Clarion follows a mandatory four-layer structure. The order is non-negotiable. AI is placed strictly downstream.
LAYER
CONTENT
RULE
1. Context
Class, subject, week, submission coverage
No AI content allowed. Must function independently.
2. Evidence
Raw signals, missing signals, conflicting signals
No summaries, no evaluative language. Evidence precedes interpretation.
3. Teacher Space
Teacher notes, tags, decisions
AI must withdraw if teacher content is present. Teacher never responds to AI.
4. AI Annotation
System annotations, insight block, silence states
Never appears first. Never concludes. Removable without breaking the system.
This is the full product structure I mapped before drawing a single screen. Four navigation destinations. Every child node governed by the four-layer rule. AI appears nowhere in the top-level structure, it is embedded, secondary, and conditional throughout.
07.1 KEY DESIGN DECISIONS
The default direction: Real-time detection. Push the moment confidence is high.
Why I rejected it: Real-time spikes anxiety, not diagnostic quality. Teachers plan weekly, the system should too.
Alternatives I considered: Real-time alerts, daily digests, event triggers. All rejected.
Tradeoff I accepted: No instant feedback.
What I gained: Lower cognitive load, calmer UX, higher signal quality, appropriate pacing.
The default direction: Require full-class uploads. Push toward 45/45 completion.
Why I rejected it: Paper classrooms can't digitize everything. 5–10 strong examples beat 45 shallow ones.
Alternatives I considered: Full-class upload, auto-digitization, student-submitted portal.
Tradeoff I accepted: Reduced coverage certainty.
What I gained: Real-world viability, adoption in paper classrooms, ethical defensibility.
The default direction: Write careful copy. Add disclaimers. Trust the words.
Why I rejected it: Copy gets edited. Structure doesn't. Restraint needs to survive a refactor.
Tradeoff I accepted: Harder to build, less immediately impressive in a demo.
What I gained: Governance that survives UI changes.
The default direction: Auto-approve above a confidence threshold. Let teachers opt out.
Why I rejected it: Silence isn't consent. Approval is an authority transfer, it has to be active.
Alternatives I considered: Confidence-based auto-unlock, approval by inaction (timeout), auto-approve with opt-out.
Tradeoff I accepted: Slower time-to-value.
What I gained: Teacher authority preserved, automation bias prevented, ethical defensibility.
The default direction: Surface more. Fill the screen. Silence reads as a bug to most teams.
Why this is the call AI couldn't make: No prompt produces this. It requires trusting restraint over noise, before any tool is open.
Alternatives I considered: Showing low-confidence insights with heavy disclaimers, showing partial insights to fill the screen, "check back later" prompts.
Tradeoff I accepted: Product feels quiet to the unfamiliar eye.
What I gained: Epistemic integrity, long-term trust, prevention of automation bias.
Hypotheses based on this week's opted-in work. You decide what gets approved.
Diagnostic Snapshot
Confidence level. Draft badge. Observation-based heading. "Written explanations show range of detail", not "Students are struggling."
Reviewed This Week
Approved and rejected insights, both visible. The system's track record, not just its output.
Action Readiness Panel
What the system chose not to show and why. This is the most trust-building element on the screen.

Insight Detail View
Nothing downstream happens without passing through here.
Confidence before Evidence seen first, on purpose. Prevents anchoring on one vivid example.
Evidence Snippets
Source context. Excerpt. Highlighted pattern. Aggregation count. No grades. No corrective language. "This is what the system noticed", not "this is wrong."
Three controls only Approve. Edit. Reject. No passive dismissal, one action is required.

Assignments Screen
Inclusion requires an explicit toggle, never assumed. "Inclusion is off by default. You choose what to include. Uploading an assignment does not trigger any analysis."
Upload ≠ Analysis, stated, not implied.

Context, Not Assessment
Empty by default. Activates only after one insight is approved. Three columns. Nothing else.
Name · Context · Submission count.
No grades. No rankings. No color coding. No flags. "Appears in 2 approved insights", factual, backward-looking, not a judgment.

Teacher-Authored Thinking Space
Activates only after approval. Notes and actions — teacher-authored, every word.
No auto-generated lesson plans. No suggested activities. No AI-drafted content. You did the diagnostic work. This part is yours.
09.1 WHAT I NEVER BUILT
These are permanent exclusions. Not “not in MVP.” Not “maybe later.” Never. I documented this list before shipping.
FEATURE
WHY IT WILL NEVER EXIST
AI detection / plagiarism flags
Accusatory, unverifiable, damages teacher-student trust
Student risk scores
Labels learners, creates stigma, surveillance dynamic
Trend graphs and trajectories
Implies evaluation, creates false momentum narrative
Auto-generated lesson plans
Reduces teacher professional identity to execution
Real-time alerts
Anxiety-inducing, breaks weekly diagnostic rhythm
Parent dashboards
Extends surveillance beyond classroom consent boundary
Predictive performance scoring
False certainty, ethical risk, no defensible evidence base
10.1 OUTCOMES
01 / Orientation
Weekly diagnostic orientation: under 2 minutes.
03 / Auto approval
Auto-approved insights: 0.
02 / Time reclaimed
Manual diagnosis time reclaimed: 3–6 hours per teacher per week.
04 / Ownership
Every decision owned by the teacher, enforced structurally.
“Clarion becomes more valuable as it becomes less assertive.”
The best outcome is a teacher who closes Clarion on Monday morning, walks into class with a clear hypothesis about where learning broke, and knows exactly what evidence supports it. That is what I designed for.
11.1 WHAT I’D DO DIFFERENTLY — AND WHAT’S STILL OPEN
01 / CONFIDENCE LAYER
Confidence needs to be visible, not implied. Show why Clarion believes something.
02 / LONGITUDINAL VIEW
The next step is weekly class insight — occurrence only, never prediction.
03 / ROADMAP RESTRAINT
The second roadmap cycle is risky. Restraint needs written rules before expansion.
04 / ETHICAL PRESSURE
The unresolved tension: what happens when institutional pressure asks for teacher-only data?
NEXT CASE STUDY:



















