Why Schools Need to Ask Which AI Model Powers Their EdTech Tools
“The product doesn’t teach your child; the model underneath does.”
Angela Chen, Innovation & Entrepreneurship Instructor at Stanford University
When a school district signs a contract with an AI edtech vendor, it is deciding which technology will sit in a room with its students, shape how they receive feedback, and influence how they think. Most procurement processes treat that decision as a privacy question, and the more consequential question, which AI model is actually inside the tool, rarely appears on the checklist. Schools, teachers, and parents are left without the information they need to make decisions for their children.
In a review by Angela Chen, an innovation and entrepreneurship instructor at Stanford University, of 50 AI edtech tools in U.S. K-12, it was found that four out of five had never disclosed which model they run. Four out of five. Schools evaluating those tools have access to privacy policies, FERPA compliance documentation, security certifications, and acceptable use agreements. What they do not have is an answer to the most basic question about what happens when a student opens the app: which AI is in the room with the child, who built it, and what values govern how it behaves.
AI tools in K-12 classrooms are no longer limited to teacher-facing planning software. They operate as tutors, writing assistants, reading coaches, and conversational learning companions. But a model built to please rather than challenge does something to a child’s intellectual development that no privacy policy was designed to prevent. A model that has never been safety-rated for minors introduces risks that FERPA compliance was never built to catch.
To better understand the disconnect between how the frameworks schools rely on were built to protect student records and not designed to evaluate what the AI is teaching, we spoke with Angela Chen, Innovation & Entrepreneurship Instructor at Stanford University.
Meet the Expert: Angela Chen, Innovation & Entrepreneurship Instructor at Stanford University

Angela Chen is an educator, advisor, and builder working at the frontier of AI and education. She teaches entrepreneurship at Stanford, where she previously grew the Stanford Accelerator for Learning and built its Education Entrepreneurship Hub from a single sentence into a program supporting 200+ scholars and 70+ ventures. Through her AI education studio, We Launch Me, she advises schools on responsible AI adoption, coaches early-stage founders, and leads live “Build with AI” sessions for non-technical professionals.
A former strategy consultant, edtech founder, and co-author of international research on AI’s impact on the future of learning and work, she serves on the AI Advisory Committee of University of Toronto Schools, is certified in Anthropic’s AI Fluency Framework, and is a Catalyst Fellow with the EDSAFE AI Alliance. She writes about the future of learning at newcurrent.substack.com.
What Schools Are Missing When They Evaluate EdTech
The standard procurement process for AI edtech is thorough in the ways it was designed to be, meaning school districts have learned to separate consumer-grade AI from purpose-built edtech.
A purpose-built tool is designed with educational context in mind: age-appropriate interfaces, curriculum alignment, and usage controls that a general-purpose chatbot does not offer. That line is real and worth drawing, and it is also not enough.
“They’re judging the packaging and ignoring the contents,” says Angela Chen, whose work at Stanford and across education entrepreneurship focuses on thoughtful AI adoption, equity, access, and K-12 innovation. “Schools have learned to separate consumer-grade AI from purpose-built edtech—a necessary line, but not a sufficient one, because ‘purpose-built’ describes the product, not the model inside it. And the model underneath is what’s actually talking to the student, shaping the feedback, and generating the lesson.”
The model underneath is also where safety ratings live. Common Sense Media rates ChatGPT-5 and Gemini K-12 as high risk for minors and rates Claude as moderate. Those grades are assigned at the model layer and do not carry through to the product built on top. A tool that passes a district’s internal review may run on a model that an independent safety organization has flagged for the exact population the district is trying to protect, and under current procurement norms, there is no mechanism requiring that information to surface.
The privacy and compliance stack that districts rely on was built to protect student records, and it does that well. What it was not built to assess is whether the underlying model has been rated for minors, how it responds to a student in crisis, and whether a mid-contract model change quietly rewrites the behavior a district originally signed off on. Model swaps are routine. A vendor may move to a new underlying model after deployment to gain capability or cut costs, and no existing framework requires them to notify the district that the engine changed. In practice, this means the district approved a behavior profile that may no longer exist.
Why the Model Underneath Is the Safety Question
Now, when a student asks for help with an essay, works through a reading exercise alone, or uses a tutoring tool late at night with no adult present, the behavior of the model running underneath the product is critical. In practice, model behavior shapes the feedback, determines how the tool responds under pressure, and governs what happens when the conversation moves somewhere serious.
“The product doesn’t teach your child, the model underneath does,” Chen says. “A tool often inherits its model’s instincts: how it behaves under pressure, and whether it indulges a student or challenges one. A sycophantic model—one built to please rather than push—trains students to outsource their own thinking to a machine that always agrees, and that tendency starts in the base model, not the layer built on top of it. The younger the student, the deeper it sets.”
With this in mind, it’s important to note that a polished interface and a well-designed curriculum wrapper tell a school nothing about how the underlying model responds when a student discloses something serious.
These student-facing tools need documented crisis off-ramps: what does the tool do when a student discloses self-harm or abuse? Does it route to a human, whether a teacher or parent, who can see the conversation, or does it remain a black box between child and model? These commitments are based at the model and policy layer, and right now most vendors are not required to document them.
The disclosure gap also carries an equity dimension that rarely gets named directly. Large districts have procurement infrastructure, legal review, and IT departments capable of asking hard questions. Meanwhile, independent schools frequently buy on relationships: a teacher loves a tool, the head of school signs off, and nothing exists between the demo and the contract but someone remembering to ask. When a parent uses public education savings account funds to buy an AI subscription directly, no institutional check exists at all.
“The burden lands heaviest exactly where the capacity to carry it is thinnest,” Chen explains, and the Education Savings Account (ESA) case is where that problem is growing fastest with the least scrutiny.
What Transparency Actually Requires
Vendor transparency on AI models does not require disclosing proprietary architecture or model weights. It requires naming what is inside the product. That disclosure costs a vendor nothing except candor, and its absence tells schools something worth knowing.
“Naming which model you run isn’t the same as exposing how you built the product,” Chen notes. “It’s a disclosure that costs a vendor nothing but candor. This should be public information, not locked behind a procurement Non-Disclosure Agreement (NDA). Schools hand over the tools, but parents absorb the consequences. Both should be able to read the answer.”
Beyond model disclosure, the baseline for any student-facing tool covers whether the underlying model has been safety-rated for minors, what behavioral commitments govern student interactions beyond the model’s defaults, whether student data trains the model, and whether the vendor notifies districts when they change or add a model. Content filters are not the same as documented commitments. A vendor may have added safety layers, removed constraints, or left the model’s defaults entirely untouched, and from the outside there is no way to tell which.
“‘We can’t describe how our tool behaves with a child because that’s proprietary’ is not a confidentiality claim,” Chen says. “It’s a refusal to answer the only question that matters. And that refusal is itself the signal: a vendor’s willingness to name what’s under the hood is never a neutral data point.”
And even more importantly, Chen explains it is shocking “how few adults are left standing between the child and the model as you move down the chain of buyers – and it thins out exactly where the child is youngest and most exposed.”
In states running universal education savings account programs, AI subscriptions fall below automatic approval thresholds and are audited only after the fact, if at all. States decide whether an AI tool qualifies as an eligible ESA expense, vendor by vendor, with no national standard and no requirement that anyone verify which model powers the product. The disclosure has to come from the vendor volunteering it and the program requiring it, because it cannot come from the parent. But right now, it seems neither is happening.
How Districts Can Build Transparency Into Procurement
Ultimately, adding model transparency to procurement does not require rebuilding the process or turning AI adoption into an evaluation gauntlet. A single checkpoint that every tool must clear creates the bottleneck that makes procurement feel like an obstruction and gives well-resourced vendors a single point to game. Neither outcome serves students.
“The rubber stamp and the bottleneck are the same mistake in different fonts,” Chen points out. “One credential that either waves every tool through or blocks them all. You avoid both by refusing to make any single credential the gate, instead relying on convergence across them as the signal.”
Effective procurement requires several inputs operating simultaneously: evidence the tool works, an independent safety rating on the underlying model, and disclosure of which model it runs. Efficacy data does not reveal what model is inside. A safety rating grades the model but not the product wrapped around it. The convergence does the work that no single gauge can do alone.
For large districts, this means adding one substantive question to an existing evaluation process and building a rubric around it. For independent schools without formal procurement infrastructure, it means one person putting the question directly to the vendor before signing. The authority is already there, and what is missing is the knowledge that the question exists to be asked.
The ESA case is the hardest, and Chen is direct about why. “That’s the gap growing fastest with the least scrutiny, and it’s the one I’d put in front of readers first.”
When public funds flow directly to families to purchase AI tools, the institutional layer disappears. The parent becomes the only check, and asking which model powers a product, whether it has been safety-rated for minors, and how it behaves with a child in distress requires technical fluency that is not evenly distributed and has nothing to do with how much a parent loves their child.
Arizona’s universal ESA auto-approves purchases under $2,000, and an AI subscription sits well within that threshold. No one at any level is required to ask which model is inside. That fluency has to be built into the system upstream, through vendor disclosure requirements and program eligibility standards, before it reaches the family.
At the end of the day, the question of which AI is teaching a child is not a technical detail to resolve after procurement. It is the procurement question, and the schools, districts, and program administrators that start asking it now are better positioned than those waiting for a framework to require it.
