STEP 01
Question and recording
A selected interview question creates an attempt with the original audio kept as the reviewable source.
A practice system that separates improvement from scoring noise
Can a daily interview-practice loop show whether an answer is actually improving, rather than merely receiving a different score?
THE PROBLEM
Interview practice usually produces a vague feeling: an answer either felt better or it did not. That is not enough to tell whether a habit is improving, whether the question changed, or whether the feedback changed.
Automated scoring adds another problem. If the scoring instrument drifts, its changing number can look exactly like progress or decline.
WHAT I DID ABOUT IT
A local interview-practice harness that records a spoken answer, keeps a verbatim transcript for delivery signals and a cleaned transcript for content review, then scores both against a fixed metric registry.
The review dashboard keeps the recording, transcript, metric evidence and human overrides together. A trends view compares repeated attempts, while a golden-set drift check reruns a scorer against labelled attempts before its results are trusted.
TRY IT YOURSELF


Demo in progress
The harness runs locally because it stores microphone recordings and interview material. The screenshots show the review and trend views from a real local run.
Want to see it sooner? Ask me.BEFORE AND AFTER
Memory of an answerAudio, transcripts and evidence
General impressionNamed metrics and overrides
One-off scoreTrends with drift checks
ARCHITECTURE
STEP 01
A selected interview question creates an attempt with the original audio kept as the reviewable source.
DETERMINISTIC
The system retains a verbatim transcript for delivery signals and creates a rule-based cleaned version for content scoring.
DETERMINISTIC
One registry defines every metric, its target, direction and weight so planning and review cannot silently disagree.
STEP 04
Independent scorers produce evidence-backed metrics, while the dashboard supports human review and overrides.
STEP 05
Repeated attempts form a trend. Re-running scorers against labelled examples detects when the instrument itself has shifted.
The important boundary is between a changing answer and a changing judge. The harness treats both as things that need evidence.
BUSINESS IMPACT
A consistent record of attempts and evidence could make interview practice more deliberate than relying on memory alone.
Separating delivery from content scoring could make it easier to identify what an answer needs to improve.
Observed means measured in a real engagement. Estimated is reasoned from the work but not measured. Potential is what the approach makes possible. Nothing here is dressed up as more than it is.
WHAT FAILED
Several imports and scorer assumptions typechecked but failed when the modules actually loaded. A scoring system that has never run is not a measuring instrument.
WHAT CHANGED
The project added import smoke tests and a coverage gate that rejects a scorer emitting an unregistered metric or a registered metric with no producer.
CODE
The implementation is a private local tool. Its architecture, verification gates and selected dashboard evidence are documented here.
View the repositorySHARE AND ENJOY
Tell me what's going on. I'll come back within a business day, and if I'm not the right person for it I'll say so.