AI engineer · 8 years · TwoWeeks

In two weeks,I turn your idea into AIyou can use right away.

I build the systems behind the feature: the data, the models, the checks and the product around them. One person takes it from the first call to operations, and every two weeks you get something you can actually use.

  • 1M+people have used AI I built
  • $500M+in yearly business my AI supports
  • 36%average performance gain in the models I built
  • 2wksfrom idea to prototype

Everyone says you need AI. So why are you still stuck?

  • 01

    No idea where to start

    You hear a lot about what AI can do, but nobody tells you concretely how it fits your own work.

  • 02

    It works in the demo, not in real use

    The chatbot that looked great in the demo gives odd answers to real customer questions and real data.

  • 03

    Outsourcing that leaves nothing behind

    When the contract ends, the code, the documents and the know-how stay with the agency.

  • 04

    Quotes with no reasoning behind them

    You get a single number, with no way to see why it costs that or what dropping a feature would save.

I only promise what works, and show it every two weeks.

  1. 01

    Tested on your data first

    I build a small version on your real documents, calls and data, and only the features that prove useful go into full development.

  2. 02

    Something that works, every two weeks

    You see progress as something you can click through, not as a status report.

  3. 03

    A quote broken down by feature

    You see what each feature costs, so you can drop what you do not need and pay only for what you do.

What I build for you.

Eight things, each one shipped before. Click a card for how it works and where it has run.

RAG · Retrieval-augmented generation

An AI assistant that understands and uses your documents

Your team asks in plain language; the assistant reads your own files, answers with the source shown, and can act on the answer — draft the reply, fill the form, open the ticket. When the source is not there, it says so.

  • RAG
  • hybrid search
  • citations
  • streaming chat
  • KakaoTalk/Slack

Running live inside the production app of a $500M+ revenue company.

How it worksClose

What you get. An ingestion pipeline for your documents (PDF, Word, Excel, slides, Korean office formats, scans), a search layer that combines meaning and keywords, and a chat interface — web or inside a messenger — that answers with citations. Access control, memory and deployment on your infrastructure are part of the build, not extras.

How I keep it honest. Before launch we agree a set of real questions and the answers a good employee would give. The assistant is measured against that set, and every later change is measured again. An answer without a source is not an answer; the system says so.

In technical terms. Retrieval-augmented generation (RAG): semantic chunking, hybrid search (vector + keyword) with reranking, citation-backed answers, streaming chat, and an evaluation set scored by an LLM judge.

GraphRAG · Knowledge graph

A knowledge graph of your expert domain

When your knowledge is about how things relate — crops and treatments, parts and failures, rules and cases — a graph answers questions a document search cannot.

  • Neo4j
  • LLM extraction
  • text-to-Cypher
  • expert review loop

Built and shipped for a Fortune 500 subsidiary — the ontology, the extraction, the database and the agent on top.

How it worksClose

What you get. A schema designed with your experts, an extraction pipeline that turns your documents into the graph with quality control (conflicts flagged and resolved against human-decided cases), the database deployed with build/snapshot/reset tooling, and an assistant that queries the graph in natural language. Method documentation so your team can extend it after handover.

Why most vendors cannot show this. The hard part is not the database — it is constructing the data with measured quality. That construction is what I have done in production.

In technical terms. GraphRAG: ontology design, LLM extraction with quality control, Neo4j multi-database, text-to-Cypher agents, and hybrid graph + vector retrieval.

AI agents · Multi-agent orchestration

Agents that research and write, with a person in the loop

A team of AI agents does the research, scoring and drafting for one business workflow; a human approves before anything leaves the building.

  • LangGraph
  • CrewAI
  • approval gates
  • resumable runs
  • cost per agent

Export-buyer reports live at FederationLabs; a 20-agent marketing pipeline with two approval gates at STIA.

How it worksClose

What you get. One workflow — market research, buyer scoring, campaign strategy, proposal drafting — designed as a pipeline of specialist agents with validation steps: grounding against real data, a critic that revises the draft, and approval gates where a person must click. Outputs land in your database as structured records, not loose text.

What I insist on. Long runs must survive a crash and resume. Approvals must be enforced in code, not in a prompt. Every model call must have a cost attached. These are the parts that make an agent system usable a year later.

In technical terms. AI agents and LLM workflows: memory, tool use, guardrails, multi-agent orchestration (LangGraph, CrewAI), deterministic pipelines for multi-step business logic and structured output, human-in-the-loop gates, resumable runs.

LLM workflows · Document generation

Finished documents, in your template

Turn source material — recordings, PDFs, web data, a curriculum — into real PowerPoint, Word and Excel files that match your templates and open ready to edit.

  • PPTX/DOCX/XLSX generation
  • OCR
  • transcription
  • schema-validated output

Classroom decks for paying teachers at MangoFactory; consulting reports from phone calls for an agricultural enterprise.

How it worksClose

What you get. Ingestion for messy inputs (scans, Korean office formats, audio), generation that fills your template in place so branding and layout survive, image selection by meaning, and delivery by e-mail, storage or inside your app. Structured output is validated against a schema before it is written, so a bad generation fails loudly instead of shipping quietly.

What it is not. Not a markdown dump pasted into a slide. The output is the same kind of file your team already makes — only faster.

In technical terms. LLM workflows with schema-validated structured output, in-place PPTX / DOCX / XLSX generation, OCR and transcription for ingestion, semantic image retrieval.

LLM fine-tuning · Model deployment · Voice AI

Your own LLM model — fine-tuned and deployed

When an API is too expensive, too slow, too inaccurate or not private enough, I fine-tune a model on your data and deploy it on your hardware — the training and the serving, both, for text and for speech.

  • LoRA / full fine-tuning
  • vLLM / ONNX serving
  • speech-to-text
  • diarization
  • GPU + CPU deployment
  • annotation tooling

In-house models replaced GPT-4 in a 1M-user product; fine-tuned speech recognition at 10.45% error on real farm calls where the commercial API scores 16.39%.

How it worksClose

Fine-tuning. A data pipeline from your documents, logs or recordings, and, where needed, an annotation tool your team uses to build the training set. Base-model selection (open-weight LLMs, speech models), LoRA or full fine-tuning on a GPU cluster, and an evaluation against the API you use today on your own test set — so the decision to switch is a number on paper, not an opinion. If the API wins, I say so.

Deployment. The model served on your infrastructure: vLLM or ONNX Runtime, GPU or CPU, containerised, behind an OpenAI-compatible endpoint so your application swaps by URL. Batch and streaming modes, autoscaling where it pays, monitoring, and a CI/CD path so the next version ships the same way. Cost per request and latency are measured before and after.

In technical terms. Fine-tuning LLMs (LoRA and full) and training proprietary models for performance, cost and privacy constraints; domain-adapted STT, diarization and streaming transcription on a GPU cluster; vLLM and ONNX Runtime serving with an OpenAI-compatible surface.

LLM evaluation · Optimisation · MCP · Agent harness

Making your AI measurable — and systematically better

“It doesn’t work well and we can’t tell why.” I turn that into numbers, then into an evaluation-and-optimisation loop that improves them release after release — the same loop that took my own models to state of the art.

  • Evaluation harnesses
  • optimisation loop
  • LLM-as-judge
  • hallucination detection
  • MCP servers
  • agent harness

A state-of-the-art scoring model built with this loop (arXiv 2309.02740); a $18K, five-star fact-checking engagement; my own unattended agent runner in daily use.

How it worksClose

Measure. An evaluation framework — criteria agreed with you, scored datasets, judge models, versioned reports — and a regression suite wired to your CI, so a model or prompt change cannot silently degrade production. Where the problem is made-up claims, a per-claim fact-checking component.

Optimise. Then the loop: error analysis on the failing cases, one change at a time (prompt, retrieval, data, fine-tuning), re-measure, keep only what moves the number, and record what did not. This is how performance improves on a schedule instead of by luck. The same loop is how I took an essay-scoring model to state of the art before it served a million learners.

Run unattended. Where the goal is autonomy: MCP servers over your data and tools, a harness with permission rules, and a scheduled runner with locking, timeouts, bounded retries, cost caps and one aggregated alert per run.

The rule behind it. Verification runs in a separate process with a different model, because a system reviewing its own work misses its own mistakes.

In technical terms. Eval harnesses, LLM-as-judge scoring, hallucination detection, error-analysis-driven optimisation; MCP servers that expose your data and APIs as agent tools; the agent harness — scheduling, retries, permissions, human-in-the-loop — so agents run unattended.

End-to-end delivery · Full stack

The whole product, not just the AI

Frontend, backend, database, deployment. No engineering team yet? I build the product around the AI as well, and hand it over in a form a team can take on later.

  • React / Next.js
  • FastAPI / NestJS
  • PostgreSQL / Supabase
  • Docker
  • AWS / GCP
  • CI/CD

MangoFactory — a paying SaaS built solo, end to end, in five months; the crop-monitoring dashboard and gateway at an LG subsidiary.

How it worksClose

What you get. A user-ready product: the app your customers log in to, the backend and database behind it, payments and exports where needed, and deployment on your own cloud — from one person who has shipped the whole thing before, so the AI does not wait for a team that is not there yet.

In technical terms. React, Next.js or Svelte frontends; FastAPI or NestJS backends; PostgreSQL, Supabase or Firestore; Docker, AWS or GCP; CI/CD and infrastructure as code from day one.

AI transformation · Consulting

Finding where AI pays off in your business

Before anything is built, we find which of your processes gain the most from AI, prove it with a small pilot on your own data, and only then scale.

  • Process diagnosis
  • two-week pilot
  • success metrics
  • roadmap

Every case study on this page started this way. The two-week prototype package is the pilot.

How it worksClose

What you get. A short diagnosis of your workflows: where time goes, what data already exists, and which steps an AI could take over with a measurable result. Then one pilot, two weeks, on your data, with a number at the end that says whether it worked. Then a roadmap that orders the rest by payoff and risk.

What it is not. Not a slide deck about “AI strategy”. The output is a working pilot and a written plan you can hand to any engineer, including one who is not me.

In technical terms. Process and data audit, candidate-workflow scoring, evaluation criteria agreed before the pilot, and a phased roadmap with the scope and quote for each phase.

From the first call to operations, built as if it were my own.

  1. STEP 1

    Requirements & data

    Together we look at the problem and the data you already have. Call recordings, documents, databases: anything works.

  2. STEP 2

    Scope & quote document

    Hours and cost per feature, plus milestones, in one document a CEO can read and decide on.

  3. STEP 3

    Two-week build cycles

    Every two weeks you get something to use, and your feedback goes into the next two weeks. Each milestone has a written test you run yourself; payment follows the test.

  4. STEP 4

    Verify, hand over, operate

    I verify performance with measured numbers and hand over all code, documents and runbooks. I stay on after launch.

AI support chatbotSample
M1 · Data cleanup · first working buildDone
M2 · Field test · accuracy workIn progress
M3 · App integration · handoverPlanned
Next demoFriday, week 2

Progress, always on one screen

  • Milestone documentWhat finishes when, on one page
  • Two-week demoA working screen, not a promise
  • Project document repoPlans, designs and decisions shared with you

Only the features you need, with a quote you can read.

  1. 1

    Hours × rate, per feature

    Every feature is broken into the hours it takes and shown in a table, so you can see what each line costs. Parts I have built before cost their integration hours, and the quote says which rows those are.

  2. 2

    Three stacking packages

    Core, then extend, then refine. You choose how far to go.

  3. 3

    Scope that fits the budget

    A smaller budget does not mean lower quality. I cut scope instead, and design it so the rest can be added later. Every scope also lists what is not included.

Ask for a quote
Sample quote (actual quotes follow a consultation)
PackageFeatureHours
1 · CoreDocument collection & cleaning pipeline24h
1 · CoreChatbot that cites its sources40h
2 · ExtendAdmin console & answer-quality dashboard32h
3 · RefineKakaoTalk & app integration24h
Start here

Two-week prototype package

Find out in two weeks, on your own data, whether the idea actually works, before you commit to a big budget.

  1. DAY 1–3

    Goals & data

    Agree on the goal and success criteria, and gather the data.

  2. DAY 4–7

    First working build

    Get one core feature actually working.

  3. DAY 8–11

    Tested on real data

    Run it on real data, measure, and fix.

  4. DAY 12–14

    Review & next scope

    Review the results together and settle the scope and quote for full development.

What you keep after two weeks
  • A prototype you can use
  • A verification report with numbers
  • Scope & quote for full development
Ask about the package

What I built, and what changed.

Eight case studies, names used with permission. Click one for the full story, including the technical detail.

  • 1M+people have used AI I built

    TestGlider, an English-test practice service. I built and served the essay and speaking scoring models.

  • $500M+in yearly business my AI supports

    Farmhannong, an LG subsidiary. My crop-consulting AI runs inside its production app and has been rolling out to all farms since May 2026.

  • 36%average performance gain in the models I built

    For example, on farm phone calls a commercial speech API got 16 of every 100 characters wrong. My model gets 10.

  • 2wksfrom idea to prototype

    A first version running on your own data, ready for you to try within two weeks.

  • 5mofrom plan to commercial launch

    MangoFactory. A paying SaaS for teachers, built alone from the first line of code to launch.

  • $5M+raised by a client on my technology

    TestGlider raised this with my scoring model as its core technology. The model was shown to be state of the art in a published paper.

The agronomy knowledge graph in Neo4j — coloured nodes for conditions, diseases, cultivation types and more, joined by relationships

Farmhannong (FarmsAll)LG subsidiary · agriculture · $500M+ revenue2025–2026

Consulting AI inside a farming enterprise's app

The company's agronomy know-how — guidelines, past reports, consultants' calls — answers farmers' questions with citations and drafts the consultants' reports for them.

  • Live in production · pilot expanding to all farms from May 2026
  • Answers cite their sources; reports cite the transcript
  • One architecture map, one gateway, sixteen repositories kept in line
Read the case study →
Bar chart of character error rate on three kinds of speech, our engine against a commercial API and the open-source base model

Farmhannong (FarmsAll)LG subsidiary · agriculture2025–2026

Speech recognition that beats the commercial API on real farm calls

Consultants' phone calls become farming diaries and consulting reports. On real farming calls our fine-tuned engine makes 10.45% character errors; the leading commercial Korean speech API makes 16.39%.

  • 10.45% error on real farming calls vs 16.39% for the leading commercial API
  • From 16.18% (open-source base) to 10.45% with ~1,300 h of curated audio + 7 h of the client's calls
  • Self-hosted on 4 × B200 · four serving modes · annotation platform in daily use
Read the case study →
TestGlider product

TestGlider (Databank)EdTech startup · nearly 1M users2020–2023

Scoring a million learners' essays and speech

An English-test practice service needed instant, reliable scores for essays and spoken answers — at a scale no human grading or per-call vendor bill could support.

  • Nearly 1M users · supported a $5M+ raise
  • In-house model replaced the LLM correction pipeline
  • Self-hosted speech-to-text replaced a per-call vendor bill
Read the case study →
Six slides from a generated classroom board-game deck, in the school's own template

MangoFactoryEdTech SaaS · paying teachers2025

Classroom-ready lesson decks, in the school's own template

Teachers describe a lesson; the system returns an editable PowerPoint in their template, plus worksheets and quizzes drawn from a 120GB+ curriculum corpus.

  • Paying users; certified teachers use it in real classrooms
  • Built solo in five months, July–November 2025
  • Fills the client's own PPTX template — nothing is rasterised
Read the case study →
First page of a generated buyer research report

FederationLabsTrade-intelligence startup2025

Export-buyer research reports, written by a team of agents

An exporter names a product; a crew of agents finds the overseas buyers, scores them on real trade records, and writes a report the sales team can act on.

  • Delivered March–June 2025, live at federationlabs.ai
  • Buyer profiles graded A/B on real trade records
  • A self-critique pass revises the report before it ships
Read the case study →
Fact-checking pipeline diagram

Upwork client (confidential)Software company · Upwork engagement2024–2025

Catching an AI's made-up claims before the user does

An AI answer is split into individual claims; each claim is checked against the reference documents, and the answer is rewritten around exactly the claims that failed.

  • 5★ review · $18K engagement
  • Shipped as a pip-installable SDK
  • Failures are per claim, so the fix is per claim
Read the case study →
Agent orchestration diagram

STIAMarketing-technology company · active engagement2025–2026

Twenty specialist agents, two human sign-offs

A marketing-campaign platform where AI drafts the strategy and the outreach, but an expert and then the client approve before anything moves — and a crashed run picks up where it stopped.

  • 20 specialist subagents with two sequential approval gates
  • Runs survive process death and resume mid-way
  • ~1,300 tests across the backend
Read the case study →
From research model to production serving

Published workApplied AI research2016–2025

State-of-the-art AI model research, put into production

I plan, run and publish model research — the essay-scoring model that reached state of the art in its paper went straight into a million-user product. The habit of measuring before shipping comes from the same place.

  • State-of-the-art result, arXiv 2309.02740 — planned, run and written by me
  • Model went straight from paper to a 1M-user product
Read the case study →

Two more projects are under NDA.

They are described below without the client, the industry or any figures, so you can judge whether the capability fits your problem.

See the confidential work

Work I cannot name.

Confidential · 2026

A conversational agent with memory and guardrails

An AI companion that remembers the user across conversations and stays inside strict behavioural limits, including when parts of the system are unavailable: it degrades gracefully instead of failing.

  • LangGraph 1.0
  • long-term memory
  • guardrails
  • evaluation harness
  • tracing
How it was built
  • Long-term memory layered over the conversation graph; guardrails include an offline language check that needs no API call.
  • An evaluation harness and full tracing, so behaviour changes are measured, not felt.
  • Graceful degradation by capability tier: the agent keeps answering with less when a dependency is down.
  • Typed, linted and tested (mypy, ruff, pytest) and handed over with documentation.
Confidential · 2026 · in production

An AI service the platform team never has to open

The AI brain and the client's platform are separated by a written interface contract with a stub mode, so their engineers integrate against a stable surface while the AI evolves behind it.

  • FastAPI async
  • job queue
  • real-time events
  • PDF/DOCX export
  • payments
  • IaC + CI
How it was built
  • Service interface specification first; a stub implementation lets the platform team build and test before the AI is finished.
  • Async API with background jobs and real-time updates; document export; payment integration; error monitoring.
  • Infrastructure as code and CI from day one; running in production.

I design it, build it, and prove it with numbers.

  • 8yrsin AI research & development
  • SOTAworld-leading results, shown in a paper
  • 100%of code and documents handed over
  • One person, end to end

    No split between sales and engineering. I take the first call, design the system, build it and run it. The person you talk to is the person who writes the code.

  • All code and documents handed over

    Code, trained models, design documents, evaluation sets and runbooks all become your assets. Vendor parts are called "integrate", never hidden as "built".

  • Verified with measured numbers

    Instead of "it works well", I report performance as numbers measured on the same data, by the same standard. Features that measure zero gain are switched off and documented.

  • Operations after launch

    Monitoring, incident response, retraining and new features: I stay on after launch.

Frequently asked questions

Q.Why two-week cycles?

With AI projects, a lot is unknown until you actually use the thing. Seeing something that works every two weeks lets us correct course early, which means less wasted money.

Q.Do we get the source code and deliverables?

Yes. Source code, trained models, data-processing scripts and design and operations documents are all handed over, and what is handed over is written into the contract.

Q.How do I pay?

Fixed price per milestone. Each milestone comes with a written test you run yourself, and payment follows the test, not the calendar. Through Upwork or a direct contract with TwoWeeks.

Q.How do we know the AI is good enough?

I check it with numbers measured on real data. For speech recognition, for example, I ran the same 2,354 call sentences through my model and a commercial service and compared the errors side by side.

Q.What if our data is a mess?

That's fine. I start by turning scattered documents, recordings and spreadsheets into something usable.

Q.Are you a freelancer or a company?

Both. I work through my company, TwoWeeks, and also through Upwork. Either way you work with me directly.

In two weeks, I'll show you something that works.

Message me on Upwork with the problem you want solved. You leave the first call with the problem stated in one sentence and a view on whether AI is the right tool for it. And if not, what is.

Working in Korea? The same work is offered through my company, TwoWeeks. twoweeks-acu.pages.dev ↗

Include these for a faster reply
  1. Your company and service
  2. The problem to solve
  3. The data you have
  4. Timeline and budget