← All work
Custom models · Replacing the LLM API · GPU serving

Scoring a million learners' essays and speech

EdTech startup · nearly 1M users · 2020–2023 · AI researcher and engineer, model to production

The in-house model did not just match the API. It replaced it in production.

An English-test practice service needed instant, reliable scores for essays and spoken answers — at a scale no human grading or per-call vendor bill could support.

Users
Nearly 1M
Funding supported
$5M+ raise
Services in production
4, with staging and production CI/CD
Research
arXiv 2309.02740 — state of the art at publication

What they needed

TOEFL and IELTS learners want to know, right now, what their essay or spoken answer would score. Human graders do not scale, and paying a vendor for every answer does not either. The scores also had to hold up: a junk answer must score zero, and the score the user sees must be the score the model produced.

What I built

The company’s own essay-scoring model — the research behind it was state of the art when published — and a speaking-scoring model that listens to the audio and reads the transcript. Later, an in-house correction model that replaced the GPT-4 pipeline, and a self-hosted speech-to-text service that replaced the vendor bill. All four ran as production services on AWS with a test harness that checked the live models on every deploy.

What changed

The product scaled to nearly a million users on models the company owned, and the AI was a core part of a $5M+ raise.

Under the hood
  • Essay scoring: transformer regression with rubric-specific heads, adversarial training and augmentation (arXiv:2309.02740). Speaking: wav2vec audio features fused with text.
  • Serving: Flask/uwsgi with Celery + RabbitMQ workers, Docker → ECR → ECS, separate staging and production pipelines. Speech-to-text ported to ONNX Runtime for cost.
  • Harness: live-model checks on deploy — junk answers must score 0, served scores must match local inference.
  • Adjacent: a TypeScript RAG application for the same product line (NestJS, hybrid vector + keyword search, React).
Model serving architecture
How the models are trained, tested against live traffic and served
PyTorchHuggingFacewav2vecONNX RuntimeFlaskCeleryRabbitMQDockerAWS ECS
testglider.com ↗ · Paper (arXiv) ↗