You open a link, a camera turns on, and a question appears on screen with a 30-second countdown before recording starts. No interviewer, no follow-up questions, no read on whether you're landing. Just you, a webcam, and a timer. This is the asynchronous AI video interview, and platforms like HireVue have made it a standard first-round filter at large employers, especially for high-volume roles in retail, finance, customer service and early-career hiring.
The format unsettles people for a reasonable reason: it feels like you're being judged by a machine on things you can't control, like your face. The good news is that most of what actually gets scored is far more mechanical, and far more within your control, than the "AI is reading your soul" anxiety suggests.
💡 Key takeaway: The score comes mostly from your words and pace, not your face. Structure your answers around what you did and said, and the algorithm has something concrete to grab onto.
What the AI Video Interview Actually Scores
HireVue itself moved away from facial analysis for scoring years ago, after research and legal scrutiny (including an Illinois biometric privacy lawsuit and pressure from the FTC) pushed the industry toward transcript-based evaluation. Most current-generation asynchronous interview tools score you primarily on what a speech-to-text engine produces from your answer, not on the video frame itself. In practice, that means three things carry the most weight:
- Keyword and concept matching: does your transcript contain the language, skills and concepts tied to the job's competency model? If the role calls for "conflict resolution" and you talk about a disagreement with a coworker without naming what you did to resolve it, the system has less to match against.
- Structure and completeness: did you answer the actual question, with a beginning, middle and end, or did you trail off, restart, or answer a different question than the one asked?
- Fluency markers: filler words like "um," "uh," and "like," long pauses, and false starts get flagged as hesitation signals. A transcript riddled with disfluencies reads to the model as lower confidence and lower preparedness, whether or not that's true.
What's mostly not in that scoring model anymore: your appearance, your race, your gender, or subtle facial "micro-expressions." That doesn't mean visuals are irrelevant to the overall impression, since a human recruiter often reviews flagged or top-scoring responses afterward, but the primary algorithmic gate is linguistic, not visual.
Why the STAR Method Wins With a Machine
Career coaches have recommended the STAR method (Situation, Task, Action, Result) for behavioral interviews for decades, mostly because it keeps humans from rambling. It turns out to matter even more for a transcript-scoring system, for a simple reason: STAR produces predictable structure and named entities. "Situation" gives the model context. "Task" gives it a goal. "Action" gives it verbs, the very words the system is often trying to match against the job's required competencies. "Result" gives it an outcome, ideally a number.
An unstructured answer like "so basically we had this really crazy situation at my old job and I kind of just figured it out and it worked out fine" gives a parsing model almost nothing to match on. A STAR-shaped version of the same story ("Our fulfillment team was missing shipping deadlines by two days. I was asked to find the bottleneck. I audited the packing queue and reassigned two staff to peak hours. Late shipments dropped by 40% within three weeks.") hands the algorithm exactly the kind of nouns, verbs and quantified outcomes it's built to reward. Practice your two or three strongest work stories in STAR shape before you ever open the recording link, so you're not building the structure live under a countdown timer.
Lighting, Shadows, and Why Your Setup Still Matters
Even with facial scoring mostly retired from mainstream platforms, some legacy systems and a handful of smaller vendors still run basic emotion or engagement detection as a secondary signal, and poor lighting is where that goes wrong most often. Harsh shadows across half your face, backlighting from a window behind you, or a dim room can all get misread by older computer-vision models as flat affect or negative emotion, simply because shadow patterns distort the facial landmarks the model is trained to read.
You don't need a ring light or a production setup. Face a window or a lamp so light falls evenly on your face, not from behind you. Sit far enough back that your face and shoulders are visible without looming into the frame. Test the recording once beforehand if the platform allows a practice question, and check that your face isn't half in shadow. This is a five-minute fix that removes a variable you don't need to be worrying about mid-answer.
Quick setup check: light source in front of you, not behind | camera at eye level | plain, uncluttered background | quiet room with no echo
Slow Down. Then Slow Down Again.
Most candidates speed up under a countdown timer, which is exactly backwards. Speaking roughly 10% slower than your natural pace does two things at once. First, it gives speech-to-text engines a cleaner signal to transcribe, which cuts down on transcription errors that can garble your keywords before the scoring model even sees them. Second, and this surprises people, it reads to human reviewers as more confident, not less. Fast speech under stress sounds like nerves. A measured pace, with brief intentional pauses instead of "um," sounds like someone who has already thought this through.
The easiest way to build this habit before the real interview: record yourself answering a practice question on your phone, then listen back with a timer. Most people are shocked at how fast they actually talk when they feel watched. Aim to leave a beat of silence between sentences instead of filling it with a filler word. A silent pause is invisible to a keyword-matching algorithm. An "um" is not.
"The most important thing in communication is hearing what isn't said." — Peter Drucker
Drucker wasn't talking about interview software, but the line applies almost too neatly here. An algorithm scoring your transcript is, in a narrow sense, listening for exactly what isn't there: the missing keyword, the answer that never names a result, the structure that never resolves. Your job in an asynchronous interview isn't to perform enthusiasm at the camera. It's to make sure the parts that matter actually make it into words.
A Practical Pre-Recording Routine
- Write out three to five STAR stories from your recent work, each with a number attached to the result
- Read the job posting once more and note the exact competency words it uses (collaboration, ownership, conflict resolution) and try to use that same language
- Test your lighting and camera angle with a mirror or a practice recording
- Do one full run-through out loud, timed, and count your own filler words
- On the real recording, pause before answering instead of starting mid-thought, then speak at a pace that feels slightly too slow to you
How Career Pilot Can Help
Career Pilot's interview preparation tools let you rehearse against realistic behavioral prompts and get feedback on structure, pacing and filler-word frequency before you're on the clock in a real HireVue-style recording. Instead of guessing whether your story landed, you get a read on whether your answer actually contains the shape (situation, action, measurable result) that both a transcript-scoring system and a human reviewer are looking for. Pairing that practice with resume and role-fit work means the story you tell in a 90-second video answer matches the story your resume already tells, which is exactly the kind of consistency both algorithmic and human reviewers respond to.
Frequently Asked Questions
Does HireVue still analyze facial expressions to score candidates?
HireVue publicly discontinued facial analysis as part of its scoring in 2021, following scrutiny including an Illinois biometric-privacy complaint and broader regulatory attention to AI hiring tools. Its current assessments are built primarily around structured interview questions and speech content rather than visual analysis. Other, smaller vendors' practices vary, so it's worth checking a specific employer's stated process if you're unsure.
How many times can I redo an answer if I mess up?
This depends entirely on the employer's configuration. Some platforms allow one practice question and a single take per real question; others allow limited retakes. Read the instructions on the login screen carefully, since this is the one detail candidates skip that changes their entire strategy.
Should I look directly at the camera or the screen?
Look at the camera lens, not the video preview of your own face, since that's what creates the impression of eye contact for anyone reviewing the recording afterward. It feels unnatural at first; practicing a few answers this way beforehand makes it feel far less strange on the real attempt.
Is it obvious if I'm reading from notes?
Often, yes. Reading word-for-word tends to flatten your pace and tone in a way both people and pacing-based algorithms pick up on. Bullet-point prompts for your STAR stories work better than a full script, since they keep the delivery closer to how you'd actually speak.
Sources and Further Reading
- Electronic Privacy Information Center, FTC Complaint Regarding HireVue (2019)
- SHRM, How AI Video Interviews Are Changing Hiring
HireVue and its competitors update their scoring models on their own release schedules, and different employers turn different modules on or off, so the mechanics described here are a snapshot of common practice rather than a spec sheet for one vendor's current build.
