← All work
Speech-to-text · ASR fine-tuning · Diarization

Speech recognition that beats the commercial API on real farm calls

LG subsidiary · agriculture · 2025–2026 · Speech stack, end to end — data, model, serving, review platform

Trained on 1% of the client's own calls. The other 99% is the roadmap.

Consultants' phone calls become farming diaries and consulting reports. On real farming calls our fine-tuned engine makes 10.45% character errors; the leading commercial Korean speech API makes 16.39%.

Farming calls
10.45% CER · commercial API 16.39% · base model 16.18%
Everyday speech
7.16% · commercial API 7.33% — nothing lost
Training data
~1,300 h curated from ~14,000 h reviewed, + 471 of the client's own calls
Serving
4 × B200, vLLM, batch / streaming / diarized

What they needed

Consultants speak with farmers by phone all day. Those calls hold the facts that should end up in the farm’s diary and the consultation report. But generic speech recognition mangled the crop names, chemicals, brand names and local terms, and users stopped trusting the transcripts. Off-the-shelf engines are trained to score well on news and studio conversation; nobody optimises them for farming calls on a mobile phone.

What I built

Four parts. A human-review platform where the team corrects transcripts and builds the training set. A fine-tuned speech model with speaker separation, served on the company’s own GPUs. A 56,000-entry farming vocabulary that feeds recognised terms back to the engine as hints. And a gateway that every other system calls — batch or streaming, with or without speaker labels. Downstream, the transcript feeds the report copilot.

What changed

Measured on 2,354 sentences from the client’s real calls, with the same scoring rules for every engine. Our engine: 10.45% character error rate. The leading commercial Korean speech API: 16.39%. The open-source model we started from: 16.18%. On everyday conversation the engine stays level with the commercial API (7.16% vs 7.33%), and on public mobile-phone recordings it is ahead (10.54% vs 11.56%). By the industry’s usual reading, around 20% is “usable but double-check” and around 10% is “trusted”. The client moved from the first band to the second — with only about 1% of their recorded calls labelled.

Under the hood
  • Data: eight public Korean speech corpora (~14,000 h) reviewed, ~11,000 h licensed and cleaned, ~1,300 h (~910k sentences) used for training with telephone-channel recordings raised to 35% of the mix — the mix, not the volume, moved the number. Telephone recordings were verified by channel fingerprint. Plus 471 of the client’s own calls (~9 h): ~7 h for training, 93 calls (~2 h) held out as the fixed judge set.
  • Acceptance rule: every candidate engine is judged on the same held-out set of the client’s calls; everyday-speech performance may not regress by more than 0.3 points.
  • Scaling evidence: with the same training volume, going from 99 to 347 client calls moved the error from 12.7% to 11.3% — roughly 0.7 points per doubling. About 1,000 h of client calls remain unlabelled; the projection is ~7% with all of them and 5% with vocabulary and post-processing work.
  • Model and serving: Qwen3-ASR fine-tuning and pyannote diarization on 4 × B200, served with vLLM behind an OpenAI-compatible /v1/audio/transcriptions surface; modes: batch, streaming, diarized, diarized-streaming.
  • Review platform: TypeScript monorepo (NestJS, Postgres, S3), 771 commits; a shared text-correction package consumed by the API, worker and web apps.
  • Measurement hygiene: same normalisation for all engines (punctuation removed, number spelling unified, spacing ignored). Inconsistent spacing in the judge set alone was worth 1.7 points, so the transcription guideline is fixed before mass labelling. A post-processing engine measured on 2,354 utterances added nothing and was documented as such.
  • Housekeeping: checkpoint purge reclaimed 9.7 TiB of GPU storage.
Character error rate on three kinds of speech — our engine, the commercial API, the open-source base
Measured 24 Sep 2026 on held-out audio, same normalisation for every engine
Voice pipeline diagram from recording to report
From recording to report — the pipeline
Audio labeler platform
The human-review platform where the training set is built
Qwen3-ASRpyannotevLLMFastAPINestJSPostgreSQLS3
Review platform (live) ↗