Skip to content

For founders — pitch, demo day, all-hands

Record three minutes. Fix one thing. Do it again tomorrow.

StageMirror counts your filler words, your pace, and every pause over 1.5 seconds, and estimates where your eyes went. Notes that can't quote you are dropped before you see them.

Free plan — 5 analyzed takes a month. No card.

REC 3:08
take_04 · 3:08 · 21 JulAnalyzed
Fillers

3.1%

target < 2%
Pace

171wpm

target 130–160
Pauses

1.0/min

target 2–3
Eye contact *

54%

target 60%–70%

Improvement

Seven filler words land in the twenty-one seconds after “the market is.” Put one full stop after “we can take a real slice of that” and let it sit.

cited 1:12–1:33 1 claim withheld
Sample analysis — illustrative, not a real user's recording.* camera-based estimate · 88% of frames read
  1. 01 · Record

    Three minutes, phone or laptop.

  2. 02 · Read

    16 fillers · 171 wpm · 3 pauses

  3. 03 · Watch

    Three passes, one channel each.

  4. 04 · Return

    Tomorrow, same time.

00:00 — the take

Three minutes is the whole ask.

Open the studio and hit record. Talk for three minutes the way you'd actually talk — no script, no retakes, no editing. Nothing is scored while you speak; live coaching would just be one more thing to perform for. Your streak advances the moment the take is uploaded, before any analysis comes back. Recording is the practice. The readout is the receipt.

REC 3:083:08

Web records at 720p today.

3:08 — the readout

Four numbers, each with a definition and a band.

A number on its own is a vibe. A number against a published target is a judgment you can act on. Three of these are counted from your audio. The fourth is estimated from your camera, and it's labelled that way everywhere it appears, including here.

Fillers

3.1%

target < 2%
How this is counted
Sixteen filler words in five hundred and twelve. Counted from a published list — um, uh, er, ah, actually, basically, literally — plus the phrases you know, I mean, kind of, sort of. like counts only when it isn't a comparison. so counts only at the start of a sentence. right counts only as a tag question. It's a word list and a set of rules. No model decides what a filler is.
Pace

171wpm

target 130–160
120130160185
How this is counted
Words divided by speaking time. Speaking time runs from your first word to your last, so the silence before you start and after you finish doesn't flatter the number. StageMirror produces one pace figure per take — not a curve. We don't draw a chart of a number we don't compute.
Pauses

1.0/min

target 2–3
How this is counted
A pause is a gap longer than 1.5 seconds between two words. You took three in three minutes. Shorter gaps exist and aren't counted — a breath isn't a pause.
Eye contact Estimate

mostly on-camera

54% · target 60%70% · 88% of frames read

Estimated from head pose and iris position at five frames a second, smoothed over one-second windows, blinks excluded. Glasses, side lighting and an off-axis camera all move this number. The angular error is roughly 5–10 degrees; what a person reads as eye contact is about 5. It is not attention, and it is not honesty. It is geometry, and it is a direction, not a score.

* Camera-based estimate.

They don't land evenly. Seven of the sixteen land in one twenty-one-second stretch.

Pauses over 1.5 seconds in the sample take
0:411.9 seconds
1:462.4 seconds
2:311.7 seconds

Three silences in three minutes. The band is two to three a minute. This is what “you never stop” looks like.

Pause positions come from word timings in your transcript.

When the camera can't read you, thin data stops being scored. If fewer than a third of your frames can be read — bad light, you walked out of shot, a covered lens — the camera estimates stop feeding the composite presence score, drop out of your progress trends, and are reported to the milestone rubrics as unmeasured rather than failed. The raw number is still shown to you, next to the coverage figure that tells you how little it rests on: you see “Only 22% of frames read — low confidence” beside it, rather than a confident-looking figure with nothing behind it.

6

Hedging words

4%

Repetition

71

Conciseness

out of 100

84 Hz

Pitch range

not monotone

68

Posture*

0.31

Sway*

11 /min

Gestures*

22%

Hands at TruthPlane*

* Camera-based estimate. Monotone is flagged when pitch variation falls below 12% of your mean.

Eight more numbers are computed and stored on every take. Some surface on the analysis and review screens, some only feed the composite scores, and some are currently just recorded — four is what a person can act on in one morning, so four is what leads.

How the scores are made

The scores are arithmetic. The language is the only part a model writes.

Clarity is the mean of a filler score and a conciseness score. Pace and pauses are band scores that taper to zero a fixed distance outside their bands. Vocal variety is a threshold on pitch variation. Presence is the mean of three camera estimates — eye contact put through the same band taper, the share of time your hands sat in the TruthPlane, and your posture score — and it's only computed at all if at least a third of your frames could be read. Overall is the mean of the spoken scores, weighted three-to-one against presence when presence exists. You can follow every one of those on paper — the 76 in the ring is what those formulas return for the numbers on this page.

The receipts

Every note points at a moment. The ones that can't are deleted.

Before any coaching reaches you, each claim's evidence is checked against your actual transcript. A quoted phrase has to be at least three words long and appear literally in what you said. A timestamp has to fall inside the take. Anything that fails is dropped — not flagged, not softened, dropped — and you're told how many were dropped.

1:12
filler: So the market is, filler: um, roughly forty billion, and we think,
1:19
filler: you know, we can take a real slice of that, filler: uh, in the
1:26
first eighteen months — filler: basically because nobody else is,
1:33
filler: like, doing the ingest layer the way we do it, so that's the wedge.

The second so, at 1:33, is not marked. The rule only counts so when it starts a sentence. like at 1:33 is marked because it isn't a comparison.

7 of 16 filler words, in 21 seconds — “you know” is two of them.

Improvement

Seven filler words land in the twenty-one seconds after “the market is.” You don't stop once in that stretch. Put one full stop after “we can take a real slice of that” and let it sit.

cited 1:12–1:33

✓ checked against your transcript

1 coaching claim could not be tied back to your transcript and was withheld.

The withheld text isn't shown to you — a claim that failed its own evidence check has no business being read. Only the count is surfaced.

Where the check stops. What's checked: strengths, improvements, camera notes, and cited moments. What isn't: the one-line summary and the single top priority, which are written from the verified numbers rather than from quotes. And a timestamp is validated as falling inside your take — not as pointing at something relevant. A quote proves the words were said, not that they were said in the context claimed. We'd rather tell you where the check stops than let you assume it goes further.

Watch it back — three times

Sound off. Picture off. Then everything.

Watching yourself once tells you nothing, because you're judging three things at once and landing on “I hate my voice.” So each pass removes a channel. Pass one is muted — you only see the body. Pass two hides the picture — you only hear the voice. Pass three gives everything back. You tick your own boxes first; the measured score stays hidden until you finish the pass, so it can't anchor you.

Pass 1 · Body

  • hands live in the TruthPlane — open palms, navel height
  • gestures are deliberate, no self-soothing
  • posture planted, shoulders level
Hidden until you finish this pass
Estimated score for this pass: 2 of 3 Estimate — pass 1's three items are all camera-derived.

Pass 2 · Voice

  • pace 130–160 wpm
  • pauses 2–3 a minute after the points that matter
  • fillers under 2%
  • the voice carries variety
Hidden until you finish this pass
Measured score for this pass: 1 of 4

Pass 3 · Full

  • the hook lands in the first 15 seconds
  • one throughline holds
  • the structure is recognisable
  • the close is decisive

Pass 3 has no tool score at all. Structure is a judgement call, not a measurement. You get the coach's narrative instead of a number we'd have had to invent.

Each pass sets up its own constraint rather than asking you to honour it: pass 1 mutes the video element and rewinds to the start when it loads, pass 2 hides the picture behind an overlay. You can still override pass 1 with the player's own volume control — the point is that the default does the work, not that we've locked you out of your own recording.

Three-pass review runs on the web app today.

Scope

We count things. We don't read minds.

Measured

  • Filler rate, filler count, word count

    lexicon and rules over your transcript.

  • Words per minute

    over speaking time only.

  • Pauses over 1.5 seconds

    counted, and placed.

  • Hedging and repetition

    counted and reported, not scored against a target.

  • Pitch range and a monotone flag

    a direction, not a verdict on your voice.

  • Estimated, from camera

    eye contact, hand position, posture, sway, gesture rate — every one labelled.

Not measured — and we won't pretend

  • Charisma.

    Nothing in the app produces that number.

  • Confidence, authenticity, executive presence.

    If a room of humans wouldn't agree on the score, a model producing one is a guess with a font.

  • Persuasiveness.

    We don't know whether your pitch worked. Neither does anything that says it does.

  • Whether you'll raise.

    No.

  • What your gaze means.

    Looking away is geometry to us, not character.

  • Whether your story is good.

    Pass 3 asks you, and doesn't answer.

An early version of the gesture counter measured 33 gestures a minute from a still photograph — pixel noise read as movement. The fix was to require movement sustained across neighbouring frames. We mention it because it's the kind of failure every camera-metric product has and almost none publish.

And then tomorrow

The point isn't today's number. It's the direction.

One take is a snapshot. Eight is a line. Progress tracks each metric against its target band, and improving means moving toward the band — so dropping from 178 words a minute to 148 counts, even though the number got smaller. Your streak advances when you record, not when the analysis lands, because the recording is the practice.

Fillers

1.6%

▼ 4.4 → 1.6 · under 2% for two takes

Pace

148 wpm

▼ 178 → 148 · in band for five takes

Streak

6-day streak

Illustrative progression. Individual results vary; we don't promise a number.

The fourteen

Two checkpoints. Twenty items. Scored on the server.

There's a Day 7 milestone rubric and a Day 14 graduation rubric. Eight items on day seven, six have to pass. Twelve on day fourteen, ten have to pass — and the bars tighten between them: fillers from under 3% to under 2%, pace from 120–170 to 130160, pauses from “at least two a minute” to “2 to 3,” eye contact from 50% to 60%. The measured items are recomputed on our server from your stored take. You supply the judgement calls. You don't supply the numbers.

  1. 1
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6
  7. 7Milestone
  8. 8
  9. 9
  10. 10
  11. 11
  12. 12
  13. 13
  14. 14Milestone

Day 7 · Milestone

8 items · pass at 6
  • Hook lands in the first 15 secondsA · S
  • One throughline holdsS · AI
  • Fillers under 3%A
  • Pace 120–170 wpmA
  • At least 2 pauses a minuteA
  • Eye contact 50% or better EstimateA
  • Hands reach the TruthPlane EstimateA
  • Confident closeS · AI

Day 14 · Graduation

12 items · pass at 10
  • Hook lands in the first 15 secondsA · S · AI
  • One throughline sustainedS · AI
  • Structure recognisable — Pixar spine, Hero's Journey or SparklineS · AI
  • Fillers under 2%A
  • Pace 130–160 wpmA
  • Pauses 2–3 a minuteA
  • Vocal variety above the monotone thresholdA
  • Eye contact 60% or better EstimateA
  • TruthPlane discipline through half your gesture time EstimateA
  • Posture 70+ with sway under 0.25 EstimateA
  • Q&A handled with a bridgeS · AI
  • Memorable closeS · AI

unmeasured2 items couldn't be measured on this take — usually thin camera data. They're reported as unmeasured rather than failed, so you're never told you did badly at something we couldn't see. Be aware of what that costs you, though: the pass bar doesn't move, so an unmeasured item still isn't a point you have. Re-record with better light and framing rather than trying to clear the bar without it.

Every automatic item is recomputed on the server from your stored analysis. The app only sends the human and AI judgements; it cannot send you a pass.
What exists today: both rubrics, and the scoring that can't be gamed. What doesn't yet: a day-by-day curriculum between them. Today you bring your own three minutes each day and the gates tell you where you stand.

Privacy

Seven specifics, and one thing that isn't built yet.

You're recording your own face. That deserves specifics, not a policy link.

What's enforced in code

  • Row-level security on every table.

    All twenty-three of them. Your rows are scoped to your user id at the database, not in application code.

  • Private storage buckets.

    Both buckets are created non-public, and every object policy requires the first path segment to be your own user id.

  • Playback links expire in five minutes.

    Three hundred seconds, one constant, used for every media URL the product hands out. There is no second, longer setting anywhere.

  • Uploads are path-checked twice.

    The server rejects any video path outside your own folder, then verifies the object actually exists before it queues anything.

  • You can't change your own plan or usage counters.

    Neither can anything running in your browser. Both are read-only to you and writable only by the service.

  • Isolation is tested, not asserted.

    A test signs up two real accounts and checks that neither can read, insert, update, upload as, or list the other. It also documents the one gap it can't close at the database layer.

  • Recordings are deleted at your retention window.

    A scheduled sweep removes the video and audio files themselves — not just the database row pointing at them — once a take is older than the window you set. Your scores and progress history stay; the footage doesn't. On the free plan two limits narrow that further: takes stop playing after 7 days until you upgrade, and nothing free is kept longer than 90 days whatever window you pick. The plan can only ever shorten a recording's life, never extend it past what you asked for.

What that promise rests on

  • “Never used to train models.”

    Worth knowing exactly what is behind it. Video frames are read on our own servers and never leave them. Audio and transcripts go to OpenAI's API, whose terms state that data submitted through the API is not used to train their models, and every call goes through a gateway we run rather than a vendor SDK — so nothing is handed to a third-party product surface by default. That is an architecture fact plus a provider's published terms, which is not the same thing as a signed, dated commitment we could show you. It is the strongest version we can honestly make, it is written up in full on our privacy page, and if it changes upstream this line changes first.

Accountability

What's shipped, and what's still being built.

Everything above this line runs today. Everything below it doesn't yet, and isn't sold anywhere else on this page.

  • Record, analyze, readoutShipped
  • Verified-citation coachingShipped
  • Three-pass reviewShipped · web
  • Progress trends, streaks, achievementsShipped
  • Day 7 and Day 14 rubricsShipped
  • Automatic deletion at your retention windowShipped
  • Day-by-day 14-day programShipped
  • Warm-up routineShipped
  • Investor and press roleplayShipped · web
  • Q&A gauntletShipped · web
  • Mobile parity for review, rubrics and practiceBuilding

Pricing

Two plans. One difference.

Free analyzes 5 takes and runs 3 practice conversations a month — enough to find out whether you'll actually keep doing this. Pro is $19 a month, or $149 a year, and takes both caps off. The 14-Day Intensive is not behind either of them.

Free

$0

  • 5 analyzed takes a month.

    Enforced on the server. The sixth is refused with a clear message, and your recording stays where it is, so an upgrade can pick it up.

  • 3 practice conversations a month.

    Investor roleplay and the Q&A gauntlet draw on the same allowance. This is the only other thing free is capped on.

  • The whole 14-Day Intensive.

    Every day and both scored gates — the programme is not a paid feature. Day 7 and Day 14 are analyzed on us even in a month where your takes are used up.

  • Everything else in the product.

    Every metric, verified-citation coaching, three-pass review, progress trends, both rubrics.

  • No card.

Record your first take

Pro

$19 /month

  • Unlimited analyzed takes and practice conversations.

    That's the difference — the two monthly caps come off, and nothing else changes.

  • Or $149 a year.

    Billed once, $79 less than paying monthly for a year.

  • Better takes, kept watchable.

    1080p instead of 720p, every recording re-watchable for its whole retention window instead of locking after 7 days, and no ads in the phone apps. The 14-Day Intensive is not on this list — it is free, and stays free.

  • Cancel in the app.

Record your first take

Start on free. Nothing on this page requires a card.

Objections

Fair questions.

Does this work with an accent, or if English isn't my first language?

Pace, pause length and filler rate are counted from a transcript, so the numbers are only as good as the transcription is — and transcription accuracy does vary by accent. The filler list is English-only and includes discourse markers like “you know” and “I mean” that are used differently across dialects. The 130–160 pace band comes from English-language public-speaking guidance and isn't tuned per speaker. You can read your own transcript, so you can check it. If the transcript is wrong, the numbers on top of it are wrong too, and you should treat the readout as broken rather than as a verdict on how you speak.

Is the eye-contact number real?

It's a real estimate and a poor certainty. It comes from head pose and iris position, carries roughly 5–10 degrees of angular error against a perceptual threshold near 5, and is degraded by glasses, low light and a camera that isn't near your eyeline. That's why it carries an ESTIMATE label everywhere it appears, why it's always shown next to the share of frames we could actually read, and why below 33% coverage it stops being scored — it drops out of your presence score, out of your progress trends, and out of the milestone rubrics, which record it as unmeasured. We still show you the number at low coverage rather than hiding it; we just stop letting it drive anything.

Can the coach make things up?

It can generate an unsupported claim — every language model can. What it can't do is show you one in the four checked sections. Each claim's evidence is matched against your transcript before storage: a quote must be three or more words and appear literally in what you said; a timestamp must fall inside the take. Failures are deleted and counted. Two caveats we'd rather state than have you discover: an in-range timestamp isn't proof the moment is relevant, and the one-line summary and the top priority aren't put through that check.

What if the analysis fails?

Then you get what did work. If the camera track can't be processed, you get the audio analysis on its own. If the coaching model is unreachable, you still get every measured number, and the screen tells you coaching is missing rather than filling the space with something generic. Nothing quietly degrades into an invented score.

What if I can't stand watching myself?

That's the normal reaction, and it's why the review is split into three passes. Start with pass two, where the picture is hidden and you only hear it. Most of what's fixable in a first week is audible, not visible. Three minutes is short on purpose, and nobody else ever sees the take.

Record your first take
StageMirror — a measurement instrument for your pitch