Skip to main content
Preterview
← All guides
Guide

Structured Interview Scorecard: Template, Evidence and Ratings

Updated · Preterview

A structured interview scorecard leaves defensible evidence when it has five parts: criteria drawn from a job analysis, the same questions asked of every candidate in the same order, a rating scale that describes observable answer levels, each interviewer's own notes, and a review step after the interviews. That sequence follows the design principles in the U.S. Office of Personnel Management's 2008 structured interview guide. A form with only score boxes turns impressions into numbers and cannot explain them later.

This guide is for hiring managers and panelists who have sat in a debrief where three interviewers disagreed by two points and none could say why. It walks through the five steps, provides a blank template with criteria, evidence, rating, and follow-up columns, and shows a before-and-after rewrite of an interviewer's note. The three-level scale is Preterview's editorial suggestion, not a validated model or any employer's official rubric.

What has to be on the form for the evidence to hold up?

The difference between a form that records impressions and one that records evidence comes down to three columns. Each criterion needs a written description of the behavior that would earn each rating, a space for what the candidate actually said, and a space for what could not be confirmed in this interview and should be checked later.

The OPM guide recommends asking applicants for the same job a consistent set of questions, applying a common rating scale, and documenting the observations behind each rating official guide. It was published in 2008, but it describes design principles rather than a dated product, so it still serves as a reference. It is not a legal requirement for private employers, and your organization's own hiring policy takes precedence.

The order matters: job analysis, then questions, then the rating scale, then individual notes, then review. Questions written without a job analysis get interpreted differently by each panelist, and notes taken without a scale never connect to a rating.

Step 1: Turn job requirements into observable criteria

Start with recurring work and less frequent but critical situations in the role. Ask people who do and supervise the work what behavior is needed, then compare that account with the job description. Replace broad labels such as “communication” with observable behavior, such as explaining a requirement change and its impact to affected teams.

  1. List recurring tasks and important exceptions in the role.
  2. For each task, write one or two sentences describing the behavior that separates good outcomes from poor ones.
  3. Keep only what an interview can observe. Certifications belong on the application review, and code quality belongs in a work sample.
  4. Select criteria you can assess within the interview time and define each one in a sentence.

If the process includes a recorded or AI-assisted round, check which criteria that method can actually assess. The role requirements should stay consistent, but each stage may gather different evidence. Candidates can use the AI interview screening guide to understand those formats.

Step 2: Writing common questions and planned follow-ups

Choose past-behavior and situational questions that fit the role. The first asks what a candidate did in an actual experience; the second asks what they would do under stated conditions. These formats provide different evidence, so choose them based on the job requirements and the experience expected of applicants.

The following is a hypothetical example for illustration. For a criterion called 'prioritizing under conflicting deadlines', the past-behavior question could be: 'Tell me about a time two deadlines overlapped and you could not finish both. What did you do first, and what happened to the other one?' The situational version could be: 'In your first week, two teams ask you for different deliverables due the same day. What do you do?'

Plan the follow-ups in advance and write them on the question sheet. If an answer drifts into what the team did, the follow-up is 'Which part did you personally handle?' If the outcome is missing, it is 'What happened after that?' Keeping follow-ups on a list means every candidate gets the same chance to fill gaps, and it stops the interview from turning into an unscripted conversation for some candidates and not others.

A Situation, Task, Action, Result note area next to each question speeds up note-taking. The MIT career center describes STAR as a way to get candidates to describe their own actions and role specifically and truthfully official guide. It is a note structure, not a scoring rule; an answer does not need all four parts to earn the top rating.

Step 3: A three-level scale that describes answers, not people

This three-level example covers one criterion: adjusting priorities when deadlines conflict. Preterview created it for illustration; it is not a validated assessment model. Adapt the anchors to the job and distinguish fluent delivery from evidence of the relevant behavior.

LevelWhat is observedNote example
3Weighs deadlines, impact and dependencies; coordinates changes and checks the outcomeNotes show the decision criteria, coordination and how completion was checked
2Gives a reason for the priority and some coordination, but evidence such as the final outcome is missingExplains delaying a lower-impact task; final result is unconfirmed
1Follow-ups still do not establish a reason for the priority or relevant coordinationSays “I did the urgent work first” without explaining the basis

If the table extends beyond the screen, swipe left or right to see all columns.

A level 1 on one criterion is not a rejection. A candidate with little work history may show something different on the situational question, so rate each criterion separately and leave the overall decision for the debrief. Before the first interview, have the panel rate the same written sample answer independently and compare; disagreements at that stage are cheaper to fix than after the candidates have been seen.

Step 4: A blank template for criteria, evidence, rating and follow-up

One row per criterion. The evidence column holds what the candidate said, as close to verbatim as your notes allow. Interpretation goes only in the rating column. The follow-up column records what this interview could not confirm and where it should be checked, such as a work sample, a second-round interview, or a reference call.

CriterionEvidence (candidate's words)RatingFollow-up
Criterion 1: ______
Criterion 2: ______
Criterion 3: ______

If the table extends beyond the screen, swipe left or right to see all columns.

Do not repeat the candidate's name, school, or contact details on the scorecard; link it by application number. How long the notes are kept and who may read them depends on your organization's data-handling policy, so check with the responsible team. This article is not legal advice.

At the top of the form, list three rules for every interviewer: keep the question order, separate observation from interpretation, and do not ask anything that is not on the sheet. Reserve a fixed slot at the end for the candidate's own questions, and note anything relevant to a criterion that comes up there. What candidates tend to ask in that slot is described in questions to ask at the end of an interview.

Before and after: rewriting an impression as an observation

The following is a hypothetical example for illustration, not a real candidate's answer. The criterion is 'prioritizing under conflicting deadlines'.

Before: 'Seems proactive and responsible. Would probably prioritize well. 3.' Another panelist reading this cannot tell what earned the rating, and there is no way to compare this candidate with the next one.

After: “When two deadlines overlapped, agreed with the manager to move one by two days and sent interim results to that task’s owner. Said the delayed task affected fewer people.” Under this example scale, record a provisional 2. Follow-up: “Final completion and impact of the delay were not established; ask about them.”

The test for a good note is whether a panelist who was not in the room could read it and arrive at the same rating. 'Seems responsible' cannot be reviewed; a record of what was said and what was missing can be discussed in the debrief exactly as written. Candidates who have rehearsed will often produce well-shaped answers, and that is fine; the note still records what was said, and the follow-up column records what was not. The candidate side of rehearsal is described in the mock interview practice guide.

Step 5: Review without averaging scores into a decision

A scoring process may use averages or totals, but the evidence still needs review. When ratings differ, compare the recorded answer with the defined anchors. A difference alone does not tell you which rater is wrong, and aggregation rules should not change from candidate to candidate.

  1. Compare the evidence and anchors for criteria where ratings differ.
  2. Ask for missing evidence. If it cannot be recovered, follow the organization’s procedure for incomplete assessment records.
  3. Hand the follow-up column to whoever runs the next stage, with the specific items to check.
  4. If a question failed to elicit anything about its criterion for most candidates, record the change to make to the question sheet.

The authorized decision-makers should use the predefined job requirements, assessment evidence and the organization’s hiring procedure. Do not turn this illustrative scale into a universal passing score. Any formal scoring policy takes precedence over the example.

Sources and scope

Key takeaways

  • Define observable criteria from the role’s recurring work and critical situations.
  • Pair each criterion with a past-behavior question, a situational question, and planned follow-ups, asked in the same order for everyone.
  • Record what the candidate said in the evidence column and keep interpretation in the rating column.
  • Review evidence and rating differences even when scores are aggregated, using the predefined hiring procedure.

Frequently asked questions

Do I need a scorecard if I am the only interviewer?

Yes. Even with one interviewer, the notes let you compare candidates against the same criteria and explain a decision later. Fill in the evidence column right after each interview, and if possible have a colleague read only the notes and assign ratings independently as a substitute for a panel debrief.

How should I record answers that go beyond the question?

Copy only the parts that relate to a criterion into the evidence column and leave the rest out or mark it as off-criteria. If time is running short, return to the scripted questions so every candidate gets the same set.

Is a three-level scale too coarse?

Use more levels if the role needs finer distinctions. What matters more than the number of levels is that each one has a behavioral description and that panelists rate the same sample answer consistently before real interviews begin.

Should follow-up questions be left to each interviewer's judgment?

Stay within the planned follow-up list where possible. If an unplanned follow-up was necessary, write down what was asked and why, and decide in the debrief whether to add it to the sheet for the next round.

Can the criteria be shared with candidates in advance?

That is a policy choice for your organization. Some employers name the criteria in the posting without sharing the questions. Whatever you decide, give every candidate the same information at the same point in the process.