Try the Sample Commands
Explore recorded sample responses in this browser demo. Choose a command or type 'help' to see what’s available.
Changing a prompt can fix one example and break another. PromptOps versions prompts, validates response shapes, and runs evaluation cases so changes have something more useful than a good first impression to go on.
The approach: Versioned prompt definitions run against evaluation cases in CI. Zod checks response structure, while regression results help identify behavior changes.
Reported project measurements: 99.4% reduction in downstream JSON parse exceptions; sub-25ms input/output assertion overhead; automated eval test suites evaluating 500+ golden cases across model upgrades.
Prompt as Code & SemVer Registry: Models prompt templates with typed input/output variables, storing versions with semantic SemVer tags (major: schema contract change, minor: prompt engineering optimization, patch: parameter/temperature tuning).
Structured Output Assertion & JSON Repair: Enforces strict Zod schema validation on model outputs, with automated markdown stripping and bracket repair for multi-provider API responses (OpenAI, Anthropic, Ollama).
Automated CI/CD Evaluation Pipeline: Executes regression test suites against prompt releases on golden datasets, computing LLM-as-a-Judge semantic similarity, exact match rates, token costs, and latency metrics.
flowchart TD
subgraph PromptRepo [Prompt as Code & Version Control]
A[Prompt Template Markdown / YAML] --> B[Zod Input/Output Schema Contract]
B --> C[SemVer Version Resolver]
end
subgraph RouterLayer [Multi-Provider Dispatcher]
C --> D[Model Dispatcher & Failover Router]
D -->|Primary| E1[Anthropic Claude 3.5 Sonnet]
D -->|Fallback| E2[OpenAI GPT-4o]
D -->|Local Sandbox| E3[Ollama Llama 3]
end
subgraph ValidationEval [Output Validation & CI/CD Eval]
E1 & E2 & E3 --> F[Zod Structured JSON Validator & Auto-Repair]
F --> G[Automated Eval Suite: Golden Datasets]
G --> H1[LLM-as-a-Judge Semantic Accuracy]
G --> H2[Latency & Token Cost Profiler]
G --> H3[Hallucination & Drift Detector]
end
subgraph ReleaseGate [Release Pipeline Gate]
H1 & H2 & H3 --> I{Eval Quality Gate: >= 98% Pass}
I -->|Pass| J[Promote to Production Registry]
I -->|Fail| K[Block CI/CD Build & Alert]
end
lib/prompt_runner.ts)
// Type-Safe Prompt Template Contract & Output Assertion
import { z } from "zod";
interface PromptDefinition<TInput extends z.ZodTypeAny, TOutput extends z.ZodTypeAny> {
id: string;
version: string;
inputSchema: TInput;
outputSchema: TOutput;
template: (inputs: z.infer<TInput>) => string;
}
class PromptRunner {
public static async execute<TIn extends z.ZodTypeAny, TOut extends z.ZodTypeAny>(
def: PromptDefinition<TIn, TOut>,
inputs: z.infer<TIn>,
dispatcher: (prompt: string) => Promise<string>
): Promise<z.infer<TOut>> {
// 1. Assert input schema conformance
const validatedInput = def.inputSchema.parse(inputs);
const promptText = def.template(validatedInput);
// 2. Dispatch to LLM provider
const rawResponse = await dispatcher(promptText);
// 3. Extract JSON and validate against output contract
const startIdx = rawResponse.indexOf("{");
const endIdx = rawResponse.lastIndexOf("}");
const cleanedJson = startIdx !== -1 && endIdx !== -1 ? rawResponse.slice(startIdx, endIdx + 1) : rawResponse;
return def.outputSchema.parse(JSON.parse(cleanedJson));
}
}
lib/evals/eval_suite.ts)
// Automated Golden Dataset Eval Suite Runner
interface EvalCase {
input: Record<string, unknown>;
expectedOutput: Record<string, unknown>;
}
async function runEvalSuite(
evalCases: EvalCase[],
runner: (input: Record<string, unknown>) => Promise<Record<string, unknown>>
): Promise<{ passRate: number; totalCases: number }> {
let passed = 0;
for (const testCase of evalCases) {
const result = await runner(testCase.input);
if (JSON.stringify(result) === JSON.stringify(testCase.expectedOutput)) {
passed++;
}
}
return {
passRate: (passed / evalCases.length) * 100,
totalCases: evalCases.length,
};
}