Scoring a million learners' essays and speech
The in-house model did not just match the API. It replaced it in production.
An English-test practice service needed instant, reliable scores for essays and spoken answers — at a scale no human grading or per-call vendor bill could support.
What they needed
TOEFL and IELTS learners want to know, right now, what their essay or spoken answer would score. Human graders do not scale, and paying a vendor for every answer does not either. The scores also had to hold up: a junk answer must score zero, and the score the user sees must be the score the model produced.
What I built
The company’s own essay-scoring model — the research behind it was state of the art when published — and a speaking-scoring model that listens to the audio and reads the transcript. Later, an in-house correction model that replaced the GPT-4 pipeline, and a self-hosted speech-to-text service that replaced the vendor bill. All four ran as production services on AWS with a test harness that checked the live models on every deploy.
What changed
The product scaled to nearly a million users on models the company owned, and the AI was a core part of a $5M+ raise.
