EVIDENCE & EVALUATION

We should be able to show that students learned something—not just that they generated something.

Splash Spectrum is being designed around measurable AI literacy behaviors: inspecting sources, checking factual claims, revising prompts, recognizing uncertainty, and explaining what the student versus AI contributed.

WHAT WE WANT TO MEASURE

Learning signals tied to the actual student workflow.

Source awareness

Can a student identify which information came from teacher materials or approved references?

Verification

Can a student check a factual claim against evidence and explain whether the evidence supports it?

AI awareness

Can a student distinguish source-backed facts, student ideas, and AI-generated creative content?

Prompt revision

Can a student change instructions intentionally and explain how the output changed?

Error detection

Can a student identify unsupported, ambiguous, or incorrect AI-generated statements?

Reflection

Can a student describe what AI did well, what required human judgment, and what they would do differently?

PILOT RESEARCH QUESTIONS

The first pilot should answer practical questions before scale.

01

Do students improve at distinguishing sourced facts from generated creative content?

Measure with defined tasks, teacher feedback, product events, and student reflection rather than vanity engagement metrics.

02

Do students become more likely to inspect evidence before accepting an AI claim?

Measure with defined tasks, teacher feedback, product events, and student reflection rather than vanity engagement metrics.

03

Can students explain how changing a prompt changes the result?

Measure with defined tasks, teacher feedback, product events, and student reflection rather than vanity engagement metrics.

04

Do teachers find the workflow instructionally useful without excessive setup or review burden?

Measure with defined tasks, teacher feedback, product events, and student reflection rather than vanity engagement metrics.

05

Are safety interventions appropriately protective without blocking normal classroom learning?

Measure with defined tasks, teacher feedback, product events, and student reflection rather than vanity engagement metrics.

06

What is the real compute and support cost per completed learning project?

Measure with defined tasks, teacher feedback, product events, and student reflection rather than vanity engagement metrics.

NORTH-STAR MEASURE

Verified Learning Projects Completed

A useful internal north-star metric is a project in which the student creates an artifact, opens the Glass Box, verifies at least one factual claim, and completes a reflection. That is intentionally more demanding than monthly active users or total generations.

CreateInspectVerify ≥1 claimReflectTeacher can review evidence
SAFETY & QUALITY EVIDENCE

Learning metrics are not enough.

Each model and release should also be evaluated for instruction following, age appropriateness, factual support, citation quality, bias, safety, latency, and cost.

Critical harmful-output escape rate
False-refusal rate
Source-support quality
Age-band appropriateness
Model-specific incident rate
Generation latency
Cost per completed project
Teacher-reported usability
EVIDENCE STATUS

No invented outcomes.

Splash Spectrum does not currently publish claims such as “students improve by X%” because those results have not yet been established through a completed study. Pilot targets may be used internally, but public efficacy claims should follow actual measurement and appropriate review.

This page describes the evaluation plan. It is not a report of completed efficacy research.
EVIDENCE PARTNERS

Interested in helping evaluate AI literacy outcomes?

Contact Splash