AI engineer · 8 years · AI agents · RAG · custom models

AI that keeps working after the demo.

I build the systems behind the feature: the data, the models, the checks, and the plumbing that lets a company rely on it.

Assistants grounded in your documents. Knowledge graphs of your expert domain. Agents with a person in the loop. Models you own. And the evaluation that proves any of it works — built, measured and handed over in a form your team can run.

1M+
learners scored by models I built and served
TestGlider
$500M+
revenue company running my AI in its production app
LG subsidiary · live
$5M+
funding round supported by the AI as the core product
Databank
5★
$18K hallucination-detection engagement, delivered as an SDK
Upwork

What I build for you

Seven things · each one shipped before

I engineer the complete path from idea to production. That means the agent workflows that ground your AI in your own data, the evaluation loops that make its behaviour a number you can track, and the full stack around it: frontend, backend, database, deployment. You get a product your customers use, not a set of components.

Now A GraphRAG consulting platform and a speech-recognition stack for an LG subsidiary · a multi-agent marketing-campaign platform, active engagement · two engagements under NDA.

01 · RAG · Retrieval-augmented generation

An AI assistant that understands and uses your documents

Your team asks in plain language; the assistant reads your own files, answers with the source shown, and can act on the answer — draft the reply, fill the form, open the ticket. When the source is not there, it says so.

Running live inside the production app of a $500M+ revenue company.

RAGhybrid searchcitationsstreaming chatKakaoTalk/Slack
How it works

What you get. An ingestion pipeline for your documents (PDF, Word, Excel, slides, Korean office formats, scans), a search layer that combines meaning and keywords, and a chat interface — web or inside a messenger — that answers with citations. Access control, memory and deployment on your infrastructure are part of the build, not extras.

How I keep it honest. Before launch we agree a set of real questions and the answers a good employee would give. The assistant is measured against that set, and every later change is measured again. An answer without a source is not an answer; the system says so.

In technical terms. Retrieval-augmented generation (RAG): semantic chunking, hybrid search (vector + keyword) with reranking, citation-backed answers, streaming chat, and an evaluation set scored by an LLM judge.

See: Consulting AI inside a farming enterprise's app

02 · GraphRAG · Knowledge graph

A knowledge graph of your expert domain

When your knowledge is about how things relate — crops and treatments, parts and failures, rules and cases — a graph answers questions a document search cannot.

Built and shipped for a Fortune 500 subsidiary — the ontology, the extraction, the database and the agent on top.

Neo4jLLM extractiontext-to-Cypherexpert review loop
How it works

What you get. A schema designed with your experts, an extraction pipeline that turns your documents into the graph with quality control (conflicts flagged and resolved against human-decided cases), the database deployed with build/snapshot/reset tooling, and an assistant that queries the graph in natural language. Method documentation so your team can extend it after handover.

Why most vendors cannot show this. The hard part is not the database — it is constructing the data with measured quality. That construction is what I have done in production.

In technical terms. GraphRAG: ontology design, LLM extraction with quality control, Neo4j multi-database, text-to-Cypher agents, and hybrid graph + vector retrieval.

See: Consulting AI inside a farming enterprise's app

03 · AI agents · Multi-agent orchestration

Agents that research and write, with a person in the loop

A team of AI agents does the research, scoring and drafting for one business workflow; a human approves before anything leaves the building.

Export-buyer reports live at FederationLabs; a 20-agent marketing pipeline with two approval gates at STIA.

LangGraphCrewAIapproval gatesresumable runscost per agent
How it works

What you get. One workflow — market research, buyer scoring, campaign strategy, proposal drafting — designed as a pipeline of specialist agents with validation steps: grounding against real data, a critic that revises the draft, and approval gates where a person must click. Outputs land in your database as structured records, not loose text.

What I insist on. Long runs must survive a crash and resume. Approvals must be enforced in code, not in a prompt. Every model call must have a cost attached. These are the parts that make an agent system usable a year later.

In technical terms. AI agents and LLM workflows: memory, tool use, guardrails, multi-agent orchestration (LangGraph, CrewAI), deterministic pipelines for multi-step business logic and structured output, human-in-the-loop gates, resumable runs.

See: Export-buyer research reports, written by a team of agents · Twenty specialist agents, two human sign-offs

04 · LLM workflows · Document generation

Finished documents, in your template

Turn source material — recordings, PDFs, web data, a curriculum — into real PowerPoint, Word and Excel files that match your templates and open ready to edit.

Classroom decks for paying teachers at MangoFactory; consulting reports from phone calls for an agricultural enterprise.

PPTX/DOCX/XLSX generationOCRtranscriptionschema-validated output
How it works

What you get. Ingestion for messy inputs (scans, Korean office formats, audio), generation that fills your template in place so branding and layout survive, image selection by meaning, and delivery by e-mail, storage or inside your app. Structured output is validated against a schema before it is written, so a bad generation fails loudly instead of shipping quietly.

What it is not. Not a markdown dump pasted into a slide. The output is the same kind of file your team already makes — only faster.

In technical terms. LLM workflows with schema-validated structured output, in-place PPTX / DOCX / XLSX generation, OCR and transcription for ingestion, semantic image retrieval.

See: Classroom-ready lesson decks, in the school's own template · Consulting AI inside a farming enterprise's app

05 · LLM fine-tuning · Model deployment · Voice AI

Your own LLM model — fine-tuned and deployed

When an API is too expensive, too slow, too inaccurate or not private enough, I fine-tune a model on your data and deploy it on your hardware — the training and the serving, both, for text and for speech.

In-house models replaced GPT-4 in a 1M-user product; fine-tuned speech recognition at 10.45% error on real farm calls where the commercial API scores 16.39%.

LoRA / full fine-tuningvLLM / ONNX servingspeech-to-textdiarizationGPU + CPU deploymentannotation tooling
How it works

Fine-tuning. A data pipeline from your documents, logs or recordings, and, where needed, an annotation tool your team uses to build the training set. Base-model selection (open-weight LLMs, speech models), LoRA or full fine-tuning on a GPU cluster, and an evaluation against the API you use today on your own test set — so the decision to switch is a number on paper, not an opinion. If the API wins, I say so.

Deployment. The model served on your infrastructure: vLLM or ONNX Runtime, GPU or CPU, containerised, behind an OpenAI-compatible endpoint so your application swaps by URL. Batch and streaming modes, autoscaling where it pays, monitoring, and a CI/CD path so the next version ships the same way. Cost per request and latency are measured before and after.

In technical terms. Fine-tuning LLMs (LoRA and full) and training proprietary models for performance, cost and privacy constraints; domain-adapted STT, diarization and streaming transcription on a GPU cluster; vLLM and ONNX Runtime serving with an OpenAI-compatible surface.

See: Scoring a million learners' essays and speech · Speech recognition that beats the commercial API on real farm calls

06 · LLM evaluation · Optimisation · MCP · Agent harness

Making your AI measurable — and systematically better

“It doesn’t work well and we can’t tell why.” I turn that into numbers, then into an evaluation-and-optimisation loop that improves them release after release — the same loop that took my own models to state of the art.

A state-of-the-art scoring model built with this loop (arXiv 2309.02740); a $18K, five-star fact-checking engagement; my own unattended agent runner in daily use.

Evaluation harnessesoptimisation loopLLM-as-judgehallucination detectionMCP serversagent harness
How it works

Measure. An evaluation framework — criteria agreed with you, scored datasets, judge models, versioned reports — and a regression suite wired to your CI, so a model or prompt change cannot silently degrade production. Where the problem is made-up claims, a per-claim fact-checking component.

Optimise. Then the loop: error analysis on the failing cases, one change at a time (prompt, retrieval, data, fine-tuning), re-measure, keep only what moves the number, and record what did not. This is how performance improves on a schedule instead of by luck. The same loop is how I took an essay-scoring model to state of the art before it served a million learners.

Run unattended. Where the goal is autonomy: MCP servers over your data and tools, a harness with permission rules, and a scheduled runner with locking, timeouts, bounded retries, cost caps and one aggregated alert per run.

The rule behind it. Verification runs in a separate process with a different model, because a system reviewing its own work misses its own mistakes.

In technical terms. Eval harnesses, LLM-as-judge scoring, hallucination detection, error-analysis-driven optimisation; MCP servers that expose your data and APIs as agent tools; the agent harness — scheduling, retries, permissions, human-in-the-loop — so agents run unattended.

See: Catching an AI's made-up claims before the user does · State-of-the-art AI model research, put into production · Twenty specialist agents, two human sign-offs

07 · End-to-end delivery · Full stack

The whole product, not just the AI

Frontend, backend, database, deployment. No engineering team yet? I build the product around the AI as well, and hand it over in a form a team can take on later.

MangoFactory — a paying SaaS built solo, end to end, in five months; the crop-monitoring dashboard and gateway at an LG subsidiary.

React / Next.jsFastAPI / NestJSPostgreSQL / SupabaseDockerAWS / GCPCI/CD
How it works

What you get. A user-ready product: the app your customers log in to, the backend and database behind it, payments and exports where needed, and deployment on your own cloud — from one person who has shipped the whole thing before, so the AI does not wait for a team that is not there yet.

In technical terms. React, Next.js or Svelte frontends; FastAPI or NestJS backends; PostgreSQL, Supabase or Firestore; Docker, AWS or GCP; CI/CD and infrastructure as code from day one.

See: Classroom-ready lesson decks, in the school's own template · Consulting AI inside a farming enterprise's app

Selected work

8 case studies · names used with permission
The agronomy knowledge graph in Neo4j — coloured nodes for conditions, diseases, cultivation types and more, joined by relationships
GraphRAG · Knowledge graph · LangGraph agent

Consulting AI inside a farming enterprise's app

LG subsidiary · agriculture · $500M+ revenue · 2025–2026

The company's agronomy know-how — guidelines, past reports, consultants' calls — answers farmers' questions with citations and drafts the consultants' reports for them.

Live in production · pilot expanding to all farms from May 2026
Bar chart of character error rate on three kinds of speech, our engine against a commercial API and the open-source base model
Speech-to-text · ASR fine-tuning · Diarization

Speech recognition that beats the commercial API on real farm calls

LG subsidiary · agriculture · 2025–2026

Consultants' phone calls become farming diaries and consulting reports. On real farming calls our fine-tuned engine makes 10.45% character errors; the leading commercial Korean speech API makes 16.39%.

10.45% error on real farming calls vs 16.39% for the leading commercial API
TestGlider product
Custom models · Replacing the LLM API · GPU serving

Scoring a million learners' essays and speech

EdTech startup · nearly 1M users · 2020–2023

An English-test practice service needed instant, reliable scores for essays and spoken answers — at a scale no human grading or per-call vendor bill could support.

Nearly 1M users · supported a $5M+ raise
Six slides from a generated classroom board-game deck, in the school's own template
LLM workflow · Document generation · SaaS

Classroom-ready lesson decks, in the school's own template

EdTech SaaS · paying teachers · 2025

Teachers describe a lesson; the system returns an editable PowerPoint in their template, plus worksheets and quizzes drawn from a 120GB+ curriculum corpus.

Paying users; certified teachers use it in real classrooms
First page of a generated buyer research report
Multi-agent · CrewAI · Research automation

Export-buyer research reports, written by a team of agents

Trade-intelligence startup · 2025

An exporter names a product; a crew of agents finds the overseas buyers, scores them on real trade records, and writes a report the sales team can act on.

Delivered March–June 2025, live at federationlabs.ai
Fact-checking pipeline diagram
Hallucination detection · LLM evaluation · SDK

Catching an AI's made-up claims before the user does

Software company · Upwork engagement · 2024–2025

An AI answer is split into individual claims; each claim is checked against the reference documents, and the answer is rewritten around exactly the claims that failed.

5★ review · $18K engagement
Agent orchestration diagram
Multi-agent · Human-in-the-loop · LangGraph

Twenty specialist agents, two human sign-offs

Marketing-technology company · active engagement · 2025–2026

A marketing-campaign platform where AI drafts the strategy and the outreach, but an expert and then the client approve before anything moves — and a crashed run picks up where it stopped.

20 specialist subagents with two sequential approval gates
From research model to production serving
State-of-the-art AI research · Model training · Evaluation

State-of-the-art AI model research, put into production

Applied AI research · 2016–2025

I plan, run and publish model research — the essay-scoring model that reached state of the art in its paper went straight into a million-user product. The habit of measuring before shipping comes from the same place.

State-of-the-art result, arXiv 2309.02740 — planned, run and written by me

Work I cannot name

Under NDA

Two recent engagements are under non-disclosure agreements. They are described here without the client, the industry or any figures — only what was built, so you can judge whether the capability fits your problem.

Confidential · 2026

A conversational agent with memory and guardrails

An AI companion that remembers the user across conversations and stays inside strict behavioural limits — including when parts of the system are unavailable, where it degrades gracefully instead of failing.

LangGraph 1.0long-term memoryguardrailsevaluation harnesstracing
How it was built
  • Long-term memory layered over the conversation graph; guardrails include an offline language check that needs no API call.
  • An evaluation harness and full tracing, so behaviour changes are measured, not felt.
  • Graceful degradation by capability tier: the agent keeps answering with less when a dependency is down.
  • Typed, linted and tested (mypy, ruff, pytest) and handed over with documentation.
Confidential · 2026 · in production

An AI service the platform team never has to open

The AI “brain” and the client's platform are separated by a written interface contract with a stub mode, so their engineers integrate against a stable surface while the AI evolves behind it.

FastAPI asyncjob queuereal-time eventsPDF/DOCX exportpaymentsIaC + CI
How it was built
  • Service interface specification first; a stub implementation lets the platform team build and test before the AI is finished.
  • Async API with background jobs and real-time updates; document export; payment integration; error monitoring.
  • Infrastructure as code and CI from day one; running in production.

How a project runs

Fixed price per milestone · you test before you pay

A call about the problem

You describe what you want to happen in the business. I ask until I can say it back in one sentence. No slides yet.

Scope and quote, line by line

Every item is a row: what it does for you, whether it already exists, hours, and price. Three stacking packs so you choose the size, not the vendor.

Milestones you can test

Each milestone has a written test you run yourself. Payment follows the test, not the calendar.

Build, with a paper trail

Decisions, deviations and progress live in a project document repository you can read at any time. Technical detail stays folded under plain-language summaries.

Acceptance pack

Each delivery comes with a report in your words: current state, what changed against the plan, what to keep in mind — and a checklist you tick off.

Handover your team can run

Code, infrastructure, evaluation sets and method notes. Vendor work is called “integrate”, never hidden as “built”.

Illustrative · what a quote looks like

Eight weeks, three packs, one table

hours × rate · figures are examples
WEEK 1
Data in, first answers
WEEK 2
Retrieval tuned
WEEK 3 · M1
You test the core
paid after
WEEK 4
Second workflow
WEEK 5
Evaluation set
WEEK 6 · M2
You test end to end
paid after
WEEK 7
Your app, your login
WEEK 8 · M3
Handover
paid after
Pack 1
Core
the one workflow that must work

Grounded answers, citations, your documents, a test set.

Pack 1 + 2 · Recommended
+ Improvement
measured, not assumed

Evaluation loop, hard cases, the second workflow.

Pack 1 + 2 + 3
+ Product
inside your app

Login, roles, deployment on your infrastructure.

ItemStatusPackHours
Document ingestion, your formatsProven112
Cited answers with refusalProven118
Your CRM as a tool the agent can useBuild216

Client documents — quote, milestone plan, milestone detail — are delivered as pages in this same design, at a private address on this site.

Working with me

Five rules I keep

Proven parts are priced, not gifted

Things I have built before cost you their integration hours, and the quote says which rows those are. You pay for the outcome, not for me relearning.

Every claim has a source

In the systems I build, a sentence without evidence is removed. In my proposals, a number without a source is not written.

Verification can only downgrade

A second pass, in a separate process, with a different model. It can flag and demote. It can never promote — so it cannot hide a failure.

Measured, then decided

Features that measure zero gain are switched off and documented, even after the work is done. The evaluation set is part of the handover.

Exclusions in writing

Every scope lists what is not included. If a request needs something I have not done — AI video, no-code automation tools — I say so before you pay.

Tools

The usual suspects
PythonPyTorchHuggingFaceLangGraphLangChainCrewAIMCPClaudeGPTGeminivLLMONNXNeo4jPostgreSQLPineconeChromaRedisFastAPIFlaskCeleryDockerAWSGCPFirestoreSupabaseReactNext.jsSvelteTypeScriptNestJS
Start with a call

Tell me what should happen in your business.

Thirty minutes. You leave with the problem stated in one sentence and a view on whether AI is the right tool for it — and if not, what is.