3.1%
For founders — pitch, demo day, all-hands
Record three minutes. Fix one thing. Do it again tomorrow.
StageMirror counts your filler words, your pace, and every pause over 1.5 seconds, and estimates where your eyes went. Notes that can't quote you are dropped before you see them.
Free plan — 5 analyzed takes a month. No card.
171wpm
1.0/min
54%
Improvement
Seven filler words land in the twenty-one seconds after “the market is.” Put one full stop after “we can take a real slice of that” and let it sit.
cited 1:12–1:33 1 claim withheld01 · Record
Three minutes, phone or laptop.
02 · Read
16 fillers · 171 wpm · 3 pauses
03 · Watch
Three passes, one channel each.
04 · Return
Tomorrow, same time.
00:00 — the take
Three minutes is the whole ask.
Open the studio and hit record. Talk for three minutes the way you'd actually talk — no script, no retakes, no editing. Nothing is scored while you speak; live coaching would just be one more thing to perform for. Your streak advances the moment the take is uploaded, before any analysis comes back. Recording is the practice. The readout is the receipt.
Web records at 720p today.
3:08 — the readout
Four numbers, each with a definition and a band.
A number on its own is a vibe. A number against a published target is a judgment you can act on. Three of these are counted from your audio. The fourth is estimated from your camera, and it's labelled that way everywhere it appears, including here.
3.1%
How this is counted
um, uh, er, ah, actually, basically, literally — plus the phrases you know, I mean, kind of, sort of. like counts only when it isn't a comparison. so counts only at the start of a sentence. right counts only as a tag question. It's a word list and a set of rules. No model decides what a filler is.171wpm
How this is counted
1.0/min
How this is counted
mostly on-camera
54% · target 60%–70% · 88% of frames read
Estimated from head pose and iris position at five frames a second, smoothed over one-second windows, blinks excluded. Glasses, side lighting and an off-axis camera all move this number. The angular error is roughly 5–10 degrees; what a person reads as eye contact is about 5. It is not attention, and it is not honesty. It is geometry, and it is a direction, not a score.
* Camera-based estimate.
They don't land evenly. Seven of the sixteen land in one twenty-one-second stretch.
| 0:41 | 1.9 seconds |
|---|---|
| 1:46 | 2.4 seconds |
| 2:31 | 1.7 seconds |
Three silences in three minutes. The band is two to three a minute. This is what “you never stop” looks like.
Pause positions come from word timings in your transcript.
* Camera-based estimate. Monotone is flagged when pitch variation falls below 12% of your mean.
Eight more numbers are computed and stored on every take. Some surface on the analysis and review screens, some only feed the composite scores, and some are currently just recorded — four is what a person can act on in one morning, so four is what leads.
How the scores are made
The scores are arithmetic. The language is the only part a model writes.
Clarity is the mean of a filler score and a conciseness score. Pace and pauses are band scores that taper to zero a fixed distance outside their bands. Vocal variety is a threshold on pitch variation. Presence is the mean of three camera estimates — eye contact put through the same band taper, the share of time your hands sat in the TruthPlane, and your posture score — and it's only computed at all if at least a third of your frames could be read. Overall is the mean of the spoken scores, weighted three-to-one against presence when presence exists. You can follow every one of those on paper — the 76 in the ring is what those formulas return for the numbers on this page.
The receipts
Every note points at a moment. The ones that can't are deleted.
Before any coaching reaches you, each claim's evidence is checked against your actual transcript. A quoted phrase has to be at least three words long and appear literally in what you said. A timestamp has to fall inside the take. Anything that fails is dropped — not flagged, not softened, dropped — and you're told how many were dropped.
- 1:12
- filler: So the market is, filler: um, roughly forty billion, and we think,
- 1:19
- filler: you know, we can take a real slice of that, filler: uh, in the
- 1:26
- first eighteen months — filler: basically because nobody else is,
- 1:33
- filler: like, doing the ingest layer the way we do it, so that's the wedge.
The second so, at 1:33, is not marked. The rule only counts so when it starts a sentence. like at 1:33 is marked because it isn't a comparison.
7 of 16 filler words, in 21 seconds — “you know” is two of them.
Seven filler words land in the twenty-one seconds after “the market is.” You don't stop once in that stretch. Put one full stop after “we can take a real slice of that” and let it sit.
cited 1:12–1:33✓ checked against your transcript
1 coaching claim could not be tied back to your transcript and was withheld.
The withheld text isn't shown to you — a claim that failed its own evidence check has no business being read. Only the count is surfaced.
Watch it back — three times
Sound off. Picture off. Then everything.
Watching yourself once tells you nothing, because you're judging three things at once and landing on “I hate my voice.” So each pass removes a channel. Pass one is muted — you only see the body. Pass two hides the picture — you only hear the voice. Pass three gives everything back. You tick your own boxes first; the measured score stays hidden until you finish the pass, so it can't anchor you.
Three-pass review runs on the web app today.
Scope
We count things. We don't read minds.
And then tomorrow
The point isn't today's number. It's the direction.
One take is a snapshot. Eight is a line. Progress tracks each metric against its target band, and improving means moving toward the band — so dropping from 178 words a minute to 148 counts, even though the number got smaller. Your streak advances when you record, not when the analysis lands, because the recording is the practice.
Fillers
1.6%
▼ 4.4 → 1.6 · under 2% for two takes
Pace
148 wpm
▼ 178 → 148 · in band for five takes
Streak
6-day streak
Illustrative progression. Individual results vary; we don't promise a number.
The fourteen
Two checkpoints. Twenty items. Scored on the server.
There's a Day 7 milestone rubric and a Day 14 graduation rubric. Eight items on day seven, six have to pass. Twelve on day fourteen, ten have to pass — and the bars tighten between them: fillers from under 3% to under 2%, pace from 120–170 to 130–160, pauses from “at least two a minute” to “2 to 3,” eye contact from 50% to 60%. The measured items are recomputed on our server from your stored take. You supply the judgement calls. You don't supply the numbers.
- 1
- 2
- 3
- 4
- 5
- 6
- 7Milestone
- 8
- 9
- 10
- 11
- 12
- 13
- 14Milestone
Privacy
Seven specifics, and one thing that isn't built yet.
You're recording your own face. That deserves specifics, not a policy link.
What's enforced in code
Row-level security on every table.
All twenty-three of them. Your rows are scoped to your user id at the database, not in application code.
Private storage buckets.
Both buckets are created non-public, and every object policy requires the first path segment to be your own user id.
Playback links expire in five minutes.
Three hundred seconds, one constant, used for every media URL the product hands out. There is no second, longer setting anywhere.
Uploads are path-checked twice.
The server rejects any video path outside your own folder, then verifies the object actually exists before it queues anything.
You can't change your own plan or usage counters.
Neither can anything running in your browser. Both are read-only to you and writable only by the service.
Isolation is tested, not asserted.
A test signs up two real accounts and checks that neither can read, insert, update, upload as, or list the other. It also documents the one gap it can't close at the database layer.
Recordings are deleted at your retention window.
A scheduled sweep removes the video and audio files themselves — not just the database row pointing at them — once a take is older than the window you set. Your scores and progress history stay; the footage doesn't. On the free plan two limits narrow that further: takes stop playing after 7 days until you upgrade, and nothing free is kept longer than 90 days whatever window you pick. The plan can only ever shorten a recording's life, never extend it past what you asked for.
What that promise rests on
“Never used to train models.”
Worth knowing exactly what is behind it. Video frames are read on our own servers and never leave them. Audio and transcripts go to OpenAI's API, whose terms state that data submitted through the API is not used to train their models, and every call goes through a gateway we run rather than a vendor SDK — so nothing is handed to a third-party product surface by default. That is an architecture fact plus a provider's published terms, which is not the same thing as a signed, dated commitment we could show you. It is the strongest version we can honestly make, it is written up in full on our privacy page, and if it changes upstream this line changes first.
Accountability
What's shipped, and what's still being built.
Everything above this line runs today. Everything below it doesn't yet, and isn't sold anywhere else on this page.
Pricing
Two plans. One difference.
Free analyzes 5 takes and runs 3 practice conversations a month — enough to find out whether you'll actually keep doing this. Pro is $19 a month, or $149 a year, and takes both caps off. The 14-Day Intensive is not behind either of them.
Free
$0
5 analyzed takes a month.
Enforced on the server. The sixth is refused with a clear message, and your recording stays where it is, so an upgrade can pick it up.
3 practice conversations a month.
Investor roleplay and the Q&A gauntlet draw on the same allowance. This is the only other thing free is capped on.
The whole 14-Day Intensive.
Every day and both scored gates — the programme is not a paid feature. Day 7 and Day 14 are analyzed on us even in a month where your takes are used up.
Everything else in the product.
Every metric, verified-citation coaching, three-pass review, progress trends, both rubrics.
No card.
Pro
$19 /month
Unlimited analyzed takes and practice conversations.
That's the difference — the two monthly caps come off, and nothing else changes.
Or $149 a year.
Billed once, $79 less than paying monthly for a year.
Better takes, kept watchable.
1080p instead of 720p, every recording re-watchable for its whole retention window instead of locking after 7 days, and no ads in the phone apps. The 14-Day Intensive is not on this list — it is free, and stays free.
Cancel in the app.
Start on free. Nothing on this page requires a card.
Objections
Fair questions.
Does this work with an accent, or if English isn't my first language?
Pace, pause length and filler rate are counted from a transcript, so the numbers are only as good as the transcription is — and transcription accuracy does vary by accent. The filler list is English-only and includes discourse markers like “you know” and “I mean” that are used differently across dialects. The 130–160 pace band comes from English-language public-speaking guidance and isn't tuned per speaker. You can read your own transcript, so you can check it. If the transcript is wrong, the numbers on top of it are wrong too, and you should treat the readout as broken rather than as a verdict on how you speak.
Is the eye-contact number real?
It's a real estimate and a poor certainty. It comes from head pose and iris position, carries roughly 5–10 degrees of angular error against a perceptual threshold near 5, and is degraded by glasses, low light and a camera that isn't near your eyeline. That's why it carries an ESTIMATE label everywhere it appears, why it's always shown next to the share of frames we could actually read, and why below 33% coverage it stops being scored — it drops out of your presence score, out of your progress trends, and out of the milestone rubrics, which record it as unmeasured. We still show you the number at low coverage rather than hiding it; we just stop letting it drive anything.
Can the coach make things up?
It can generate an unsupported claim — every language model can. What it can't do is show you one in the four checked sections. Each claim's evidence is matched against your transcript before storage: a quote must be three or more words and appear literally in what you said; a timestamp must fall inside the take. Failures are deleted and counted. Two caveats we'd rather state than have you discover: an in-range timestamp isn't proof the moment is relevant, and the one-line summary and the top priority aren't put through that check.
What if the analysis fails?
Then you get what did work. If the camera track can't be processed, you get the audio analysis on its own. If the coaching model is unreachable, you still get every measured number, and the screen tells you coaching is missing rather than filling the space with something generic. Nothing quietly degrades into an invented score.
What if I can't stand watching myself?
That's the normal reaction, and it's why the review is split into three passes. Start with pass two, where the picture is hidden and you only hear it. Most of what's fixable in a first week is audible, not visible. Three minutes is short on purpose, and nobody else ever sees the take.