Aria Han

Bootstrap loop shipped

model-familiarity-engine

I wanted to know what a model had earned across a working relationship, not where it ranked.

The problem

Single-shot benchmarks don't capture how a model behaves across a real working relationship.

What I built

Onboards language models by simulating real user conversations, then builds evidence-backed model cards from observations instead of a ranking. All benchmarks are drawn from real conversation transcripts.

The replay-bootstrap loop is shipped: known-outcome tasks, redaction, replay, model cards built from what was actually observed.

Proof

Replay-bootstrap loop shipped: redaction, replay, observed model cards. MIT licensed.

What I learned

Replay with redaction lets you bootstrap evaluation from your own transcripts. No synthetic tasks required.

Designing a novel eval methodology and shipping the loop, not just the idea.

Stack

Python · Bedrock · Ollama · Claude CLI

status
Bootstrap Loop Shipped
stack
Python · Bedrock · Ollama · Claude CLI
scope
Replay · Redaction · Model Cards
license
MIT

Links

The question was never which model is best. It's what this one has earned.

Connected work