Bootstrap loop shipped
model-familiarity-engine
I wanted to know what a model had earned across a working relationship, not where it ranked.
The problem
Single-shot benchmarks don't capture how a model behaves across a real working relationship.
What I built
Onboards language models by simulating real user conversations, then builds evidence-backed model cards from observations instead of a ranking. All benchmarks are drawn from real conversation transcripts.
The replay-bootstrap loop is shipped: known-outcome tasks, redaction, replay, model cards built from what was actually observed.
Proof
Replay-bootstrap loop shipped: redaction, replay, observed model cards. MIT licensed.
What I learned
Replay with redaction lets you bootstrap evaluation from your own transcripts. No synthetic tasks required.
Designing a novel eval methodology and shipping the loop, not just the idea.
Stack
Python · Bedrock · Ollama · Claude CLI
- status
- Bootstrap Loop Shipped
- stack
- Python · Bedrock · Ollama · Claude CLI
- scope
- Replay · Redaction · Model Cards
- license
- MIT
Links
The question was never which model is best. It's what this one has earned.