How to Create a Job-Specific Technical Test for Applicants
AssessExpert Team · June 7, 2026
A job-specific technical test built from the actual job description predicts on-the-job performance dramatically better than a generic skill test. The method is mechanical — extract the skills from the JD, write a section per skill, calibrate against current employees, validate against actual hires. Most teams skip the calibration and validation steps, which is where most tests fail.
Step 1 — Extract the skills from the job description
Open the live job description for the role. List every skill mentioned. Sort into "must have" and "nice to have." Drop the nice-to-haves from the test — they add noise and lengthen the assessment without improving signal.
For a typical mid-level technical role, the must-have list is usually four to seven skills. Examples for common roles:
- Mid-level Python backend developer: Python fundamentals, async/concurrency, database design, API design, debugging skill.
- AutoCAD draftsman L2: drawing standards, layer discipline, dimensioning, block management, drawing setup.
- Financial analyst: spreadsheet logic, variance analysis, financial modelling, presentation of findings, communication.
- BIM coordinator: Revit modelling discipline, clash detection workflow, BCF reporting, federation management, coordination communication.
If your skill list is longer than seven items, the JD is too broad — split the role into two before testing. If it is shorter than three, the role is probably too junior for technical testing and you should hire on potential.
Step 2 — One section per must-have skill
For each must-have skill, write five to eight multiple-choice questions that test recall plus one short practical task that tests application. The MCQ section catches surface-level knowledge gaps; the practical section catches inability to apply.
The discipline this enforces matters. If you cannot write the practical for a skill, the skill is too vague to be testable — and probably too vague to be in the JD at all. "Strong communication" is a JD line that cannot be tested. "Writes technical documentation that a non-specialist can follow" is a JD line that can.
Time-box each section. 10 minutes per MCQ section, 15-20 minutes per practical, capped at 90 minutes total. A test longer than 90 minutes loses candidates to fatigue and drops completion rate below the threshold where the data is useful.
Step 3 — Calibrate against current employees
This is the step most teams skip and most tests fail because of. Have two or three current team members at the target role level take the test cold. They should not know the questions in advance, and they should know their results will not affect their employment.
If they score below 80%, the test is too hard or written badly. The questions are testing trivia rather than skill, or the practical is unrealistic for the time limit, or the rubric is mis-calibrated. Revise until current top performers consistently pass at the expected level.
If they score 100%, the test is too easy. It will not discriminate between candidates. Add harder questions or tighten the rubric.
If they disagree on what the right answer is for an MCQ, the question is ambiguous. Rewrite it.
The calibration loop usually takes one to two passes. The first pass surfaces obvious mis-calibrations; the second pass tunes the difficulty. Skipping this step is the single most common cause of test failure — the test goes live, rejects good candidates, and the hiring manager loses faith in the score.
Step 4 — Validate against actual hires
The calibration step gets the test working. The validation step proves it predicts performance. After three months of live use, look at the people you hired who passed the test. Were the high scorers also strong on the job? Were the low scorers struggling? If not, the test is measuring the wrong thing — and you have time to fix it before too many decisions ride on it.
The validation method:
- For each hire who has been in role for at least 90 days, rate their on-the-job performance on a 1-5 scale.
- Pull their assessment score.
- Plot the two. There should be a positive correlation.
If the correlation is weak or absent, the test is not predictive. Time to redesign — usually by tightening the practical section or replacing weak MCQs. If the correlation is strong, the test is doing its job and the calibration was right.
Common pitfalls to avoid
Trivia questions. "Which year was Python first released?" tests memory, not skill. Drop these. Test what the candidate needs to do, not what they need to know about the history of the tool.
Overly clever questions. Questions designed to catch out candidates who do not read carefully test reading skill rather than the named skill. Use them sparingly.
Practical tasks with unrealistic constraints. A practical that takes a senior engineer 90 minutes will take a strong candidate at the target level 60 minutes and a weak candidate forever. Time-box realistically.
Rubrics that depend on interpretation. "Code is clean" is unscoreable. "Functions are under 30 lines and named meaningfully" is scoreable. Specificity beats sophistication.
Tests that don't match the real work. If your developers spend their day in CI debug sessions and your test asks them to whiteboard linked lists, you are testing the wrong thing.
How long custom test design takes
A team building a custom test from scratch typically needs:
- 1 day to extract skills and structure the test.
- 3-5 days to write questions and practical tasks.
- 2 days for calibration with current employees and revision.
- 3 months of live use plus a half-day validation review.
Total upfront: about two weeks of focused work. The validation review is recurring — quarterly is healthy.
When to commission a custom test vs use a pre-built one
Pre-built tests are fine when your role looks like a standard role in the market. Mid-level Python developer, junior AutoCAD draftsman, mid-level financial analyst — these are common enough that calibrated banks exist and work well.
Custom tests are needed when your role involves proprietary tools, internal workflows, regulatory frameworks unique to your industry, or an unusual combination of skills. The off-the-shelf test will under-discriminate, and you will hire on the assessment data without actually knowing whether the candidate can do the job.
How AssessExpert supports custom test design
Our Exam Setup team builds role-specific banks to your spec — 500 questions per assessment type — and calibrates the practical task against your evaluation rubric. The build process takes two to three weeks and requires four to six hours of SME time. The bank stays private to your workspace; we do not share custom questions across clients.
See Custom Assessment Tests for the full build process. For role-specific examples in engineering, see CAD, BIM and Engineering Assessments.
FAQ
How long should building a custom test take?
About two weeks of focused work for a single role, assuming a subject matter expert is available for four to six hours across the build.
How often should the test be revised?
Quarterly review for high-volume roles, annually for low-volume. Skills required for a role drift over time and the test should drift with them.
Can the same test work for L1 and L2 versions of the same role?
Usually no. The pass mark differs, but more importantly the practical task should differ. L1 and L2 should be separate tests sharing some structural similarity.
What if our top employees can't pass the test?
The test is mis-calibrated. Revise the questions and practical until top current performers consistently pass at the expected level. This is a feature of the calibration step, not a problem.
Do candidates ever complain that the test is unfair?
Occasionally. The fix is rubric transparency — publish the rubric internally so feedback is anchored to specifics, and explain to declined candidates which sections they scored below threshold on.
What about hiring for roles that don't exist yet?
Test for the closest existing role and treat the assessment as a directional signal, not a verdict. Pure greenfield roles are usually hired on potential plus interview, with assessment as a smaller weight.
Next steps
If you want a custom test built for a specific role, the first call is a 30-minute conversation about the role and the must-have skills. Book a demo and our Exam Setup team will scope the build for you.