Notes from the review loop.
How we think about grading AI instructions against a rubric, writing rewrites you can test, and versioning prompts, agents, and skills like the production code they are.
Looking for a PromptLayer alternative? Check which job you are replacing
PromptLayer is four products in one: a prompt registry, request logging, output evals, and workflows. Most people searching for an alternative want to swap out one of them. Which one decides whether you should switch at all.
Looking for a promptfoo alternative? Decide what you are testing first
Most people searching for a promptfoo alternative do not have a promptfoo problem — they have a test-case problem. An honest read on what promptfoo is excellent at, the four reasons teams go looking anyway, and where a rubric-first tool actually belongs.
How to score a prompt from the command line
Grade an instruction, read the findings, and fail a build on the score without leaving your terminal. Here is the local loop — and why the scoring script you were about to write yourself goes stale.
What is CRIT in AI? The Socratic scoring template, and where it stops
CRIT scores a document by decomposing it into a claim and its reasons, then rating each argument for validity and each source for credibility. How it works, what it is genuinely good at, and why it does not grade your prompt.
How to test a prompt: a practical guide to prompt regression testing
A working procedure for testing prompts the way you test code — a golden set, pass criteria you can defend, and a CI gate that catches the regression before your users do.
LLM evaluation metrics: what to measure and when each one lies to you
Exact match, BLEU, embedding similarity, LLM-as-a-judge, rubric scores. Each measures something real and each fails in a specific, predictable way. Here is which to reach for and what it hides.
How to get statistical significance for your prompts
The honest answer has two halves: what real significance costs, and what you can get for far less. Both start by measuring how much your evaluator moves when nothing changes.
The ten dimensions we grade every instruction on
One number is easy to argue with. Ten named dimensions are not. Here is the rubric behind every Rubrkit score, and what each one is actually checking for.
Ship the proof, not the promise
"I improved the prompt" is a claim. A proof report is evidence — the same instruction, before and after, with the eval it now passes and the version hashes to prove which is which.
How we grade our skills with CI
We open-sourced a library of skills — then wired our own quality gate to grade them on every push. A skill that drops below the bar turns the build red.
Audit the prompt before you run the agent
A long prompt is cheapest to fix before it runs. Here is how to grade and harden an instruction over MCP, in the same client your agent lives in, before you hit go.
Stop grading prompts on vibes
A prompt that "feels good" is not a prompt you can trust. Here is why we grade instructions against a rubric instead of a gut feeling.
Anatomy of a testable rewrite
A weak instruction, the rubric dimensions it fails, the rewrite that fixes them, and the eval that proves the rewrite holds.
Version your prompts like code
Prompts, agents, and skills are production instructions. They deserve history, provenance, and a single source of truth — not a paste buried in a chat log.