← All work
State-of-the-art AI research · Model training · Evaluation

State-of-the-art AI model research, put into production

Applied AI research · 2016–2025 · Author

An experiment is only finished when it runs in a product.

I plan, run and publish model research — the essay-scoring model that reached state of the art in its paper went straight into a million-user product. The habit of measuring before shipping comes from the same place.

Paper
Rubric-Specific Approach to Automated Essay Scoring with Augmentation Training
Result
State of the art at publication
Other topics
Prompt selection with LLM feedback · RLHF · hallucination detection
Earlier
Sales forecasting · disease-surveillance clustering · smart farming

The published work

Rubric-Specific Approach to Automated Essay Scoring with Augmentation Training (arXiv:2309.02740). Planned, designed, executed and authored by me, from research question to published paper: a state-of-the-art result on automated essay scoring. The model it describes went into production and scored essays for nearly a million learners.

Other research

  • Optimised prompt selection using a task-specific model and LLM feedback (conference submission).
  • Reward modelling and reinforcement learning from human feedback (T5 reward model + PPO).
  • Hallucination detection for LLM answers (2025) — shipped as the fact-checking SDK.
  • Earlier applied work: SKU-level and agri-food sales prediction, spatial clustering of migratory birds for avian-flu surveillance, smart-livestock and soil-sensor modelling, and risk signals from PC log data.
Why this matters for a client

The research habit shows up in delivery as evaluation sets, judge models and before/after numbers — and in the willingness to switch a feature off when the measurement says it adds nothing.

PyTorchHuggingFaceT5PPO
Paper (arXiv) ↗