A self-paced course line

Ship LLM systems solo. The review is a harness of judges.

Anyone can call an LLM — that's the commodity. Turning a flaky generator into a system you can ship is review, QA and security work you'd never run by hand. This course line teaches you to run it solo — as code.

The thesis: an LLM is a commodity you'll swap in a month; the judges you build around it — grading every output against your bar and gating the release when it falls short — are what's yours. You set the standard; the harness enforces it across thousands of outputs, the review, QA and eval coverage one person could never run by hand. LLM is the commodity; the harness is the moat.

Explore the options Start free — the primer
58HTML lessons
57runnable labs
40+sources cited
8tracks

Course hero under final review

The Judge-Driven Harness · bilingual course hero

AI-generated fictional presenter

One method, eight applications

Generate → judge → gate → retry

Every track is the same spine applied to one domain. Learn it once in the anchor course, then watch it secure an app, gate a release, govern a fleet of coding agents, or grow an eval set that compounds.

generate judge (as code) gate (go / no-go) retry · promote · hold · rollback
Why this isn't a free blog post

Proof, not vibes

Three things no marketplace course bundles together.

◆ published research

Research-backed

Every claim traces to published, public research — OWASP LLM & Agentic Top 10, MITRE ATLAS, RAGAS, MT-Bench, DORA, NIST AI 600-1, C2PA — cited and credited, never hand-waved.

▶ runnable

57 deterministic labs

The 58 lessons are paired with 57 Python labs that run offline, stdlib-only, with the same output every time. Not slides — code you run, then point at your own system.

✓ honest

Shipped vs backlog

Every metric is illustrative and labelled. We teach you to measure your own numbers — never "our system hits X%." The honesty is the product.

The eight tracks

The method, applied to your domain

Start with the anchor. Add the track that matches your work — or take the whole harness.

CORE · the method

The Judge-Driven Harness

The domain-agnostic spine: discover failures, write judges as code, validate them against a compact labeled set, gate, retry, calibrate, wire into CI.

free primer + 9 modules · 8 labs
EVAL · eval

Ship LLM Apps on Evidence

Traces-first (OpenTelemetry), golden datasets, the RAG triad reference-free, readiness scoring & the Pareto frontier, reproducible CI gates, the production loop.

7 modules · 7 labs
AGENT · agentic coding

Agentic Coding Harness

Claude Code in production: AGENTS.md as contract, specialized subagents, hardening & determinism, red-team your own agents, judge-driven review — governance, not a feature tour.

7 modules · 7 labs
SEC · security

Securing LLM Apps

OWASP LLM & Agentic Top 10, MITRE ATLAS, prompt-injection taxonomy, red-team as a repeatable harness, detection-as-code → SIEM, compliance derived.

6 modules · 6 labs
DOG · dogfooding

Dogfooding

Validate your own product: run it as the eval, orthogonal quality dimensions, calibrated gates, the living question bank, and dogfooding wired into CI — ship on evidence, not "looks fine."

7 modules · 7 labs
SIM · synthetic users

Synthetic Users & Personas

Goal-directed NPC simulation, multi-turn exploration, and the honest breadth/mass validity boundary: personas generate the space, not the calibration. The limit nobody else teaches.

8 modules · 8 labs
DATA · the data layer

The Data Flywheel

The eval set is the asset, not the model: curate a seed bank, grow it from real failures, stratify & cover, dedup, judge-label with κ agreement, version and catch regressions — the bank that compounds.

7 modules · 7 labs
COST · the efficiency layer

The Cost & Latency Harness

Quality isn't free: the three-axis gate, latency percentiles over means, the cost/quality Pareto frontier, route-cheap-escalate-hard, semantic caching, and budget gates that fail regressions in CI.

7 modules · 7 labs
Launch options · individual self-study

Choose the scope that matches your work.

Core method → Complete bundle → Full Harness. Review what each option includes now; purchase links and final pricing appear only after launch approval.

Core
The anchor method — the spine every track builds on.
At launch
Pricing pending final approval
  • Free primer + 9 anchor modules
  • 8 runnable labs
  • The full generate→judge→gate→retry method
  • Free updates
Complete bundles
The anchor + one track, focused on your domain.
At launch
Pricing pending final approval
  • Anchor + one full track (or the Validate pair: dogfooding + synthetic users)
  • All that track's labs + briefs; Coding Complete pairs the AI Coding OS pack
  • À la carte — security · eval · agentic coding · dogfooding · synthetic users · data · cost
  • Free primer included · free updates

New here? The Level-0 primer is free — read it while launch details are finalized.
Want the toolkit, not the course? Pair any track with the AI Coding OS pack → (18 install-ready skills) — the harness, pre-built.

Before you buy

Straight answers

Is this beyond a free OWASP page / blog post?
Yes. The free material tells you the risks exist. This teaches you to build the harness that catches them at scale — red-team as a repeatable suite, detection-as-code, judge-driven gates — with runnable labs and paper provenance. The differentiator is the method + the code, not a list.
Are the labs real, or toy problems?
Real and runnable: 57 Python labs, offline, stdlib-only, deterministic (same output every run). Numbers are illustrative by design so you learn the pattern, then swap the single HOOK for your own model and measure your own system.
Is this current for 2026?
The content tracks the 2026 frontier: OWASP LLM (2025) + Agentic (2026) Top 10, current OpenTelemetry GenAI conventions, MT-Bench judge-bias findings, NIST AI 600-1, C2PA + SynthID, the EU AI Act Article 50 timeline. Tool status (e.g. OpenAI's announced agreement to acquire Promptfoo) is flagged with OSS alternatives.
Isn't Claude Code already taught free by Anthropic?
The free courses teach the features — subagents, MCP, skills, hooks. The Agentic Coding track teaches what they don't: governance, hardening, red-team, and quality gates — Claude Code in production, by people who ship with it — with runnable judge-labs and a starter you keep. It sits above the feature tier, not next to it.
Individual only — is there a team version?
Today every tier is individual self-study, priced accordingly. A team tier with progress tracking and seats is on the roadmap, not shipped — so we don't charge for it yet.
How do I receive it, and does it work offline?
A single .zip with an index.html launcher: open it in any browser, read on any device, run the labs on a desktop. No account, nothing calls home. Updates arrive through your store library.