Every now and again an older idea circles back around in selection psychology and finds a new audience. The current example is skills-based assessment, sometimes called “skills-first” hiring: the idea that the best way to select people is to have them complete job-style tasks, typically auto-generated from a job description and scored by artificial intelligence, and to rank them on the result. It is increasingly promoted for early-careers hiring.
The novelty is in the packaging. Strip away the interface and the algorithm and what remains is the work sample test, a method selection psychologists have been building, studying and refining for the better part of a century. The tasks are now generated automatically rather than designed by a job analyst, scored by a model rather than a trained assessor, and marketed as a break from the psychometric tradition rather than what it actually is: a well-worn branch of it, revived with the caveats left off.
To be clear about what this article is not arguing: assessing what people can do, rather than what their CV says, is a sound idea. That is precisely why work sample testing has survived for so long, and skills testing has a legitimate place in a well-designed process. The concern here is narrower and more serious. The approach is now being promoted specifically for graduate and early-careers recruitment, and that is the one population for which its entire logic breaks down. If you run a graduate programme, or advise one, this matters to you.
The graduate paradox
A work sample test rests on a simple premise: if you want to know whether someone can do the job, watch them do a piece of it. That premise holds when the applicant has done the job, or something very like it, before. It collapses when they have not.
A graduate applying for their first analyst, consulting or customer-facing role has, by definition, not yet done the work. The spreadsheet exercise, the mock stakeholder email, the ticket-triage simulation, the short business case: these are precisely the tasks the organisation will teach them in their first weeks of onboarding. Look at your own programme’s induction and you will almost certainly find it built around exactly this material. So the assessment asks candidates to demonstrate, before they are hired, the skills you are about to provide as part of the job.
What does a score on such a task actually tell you about a twenty-two-year-old? It tells you who has already been exposed to that specific kind of work: the student who landed the right internship, whose university ran the right software, whose family happened to be in the industry. It does not tell you who will be the strongest performer once everyone has been through the same training. Those are different questions, and for graduates only the second one matters.
Advocates of skills-based assessment commonly present role-specific tasks as well suited to graduates because they lack prior work experience: the argument runs that a task levels the field where a CV cannot. The opposite is true. The absence of prior experience is exactly the condition under which a work sample stops measuring capability and starts measuring exposure.
What the research says about experience and work samples
This is not a novel objection. It has been stated in the literature since the modern era of selection research began.
Schmidt and Hunter’s 1998 review, still the most cited paper in the field, was explicit that job knowledge tests and work samples are only appropriate for applicants who already know the job, and are of no use with inexperienced applicants. Roth, Bobko and McFarland, in their 2005 meta-analysis of work sample validity, made the same caveat: the good validity figures come from studies of experienced applicants. In the 2022 re-analysis by Sackett, Zhang, Berry and Lievens, which corrected the older literature for statistical over-adjustment, work samples and job knowledge tests remained respectable predictors, but nothing in that paper extends the finding to people who have never done the work.
Why does experience matter so much? Schmidt, Hunter and Outerbridge showed in 1986 that job performance is driven by job knowledge, and that job knowledge is driven by two things: general cognitive ability and time on the job. Ability determines how fast and how completely a person learns the work. Experience provides the opportunity. A job knowledge or work sample test measures the output of that learning process. For an experienced applicant the output is a fair summary of what they have acquired. For a graduate the output does not yet exist, so the sensible thing to measure is the input: the ability, and the dispositions, that determine how well they will learn.
Cognitive psychology reaches the same conclusion from a different direction. John Anderson’s work on skill acquisition distinguishes declarative knowledge (knowing that: facts, rules, procedures you can describe) from procedural knowledge (knowing how: skills that have become fluent through practice). Kanfer and Ackerman showed that early in learning any new skill, performance depends heavily on general cognitive ability, because the learner is still building the declarative foundation; only after extended practice does performance become proceduralised and depend on the specific experience gained. Skills-based tasks tap declarative and procedural knowledge of a specific job. Graduates are, by definition, at the start of that curve. The construct the approach measures is the one construct graduates cannot yet be expected to have.
What skills-based assessment cannot see
Suppose the task is well built and fairly scored. There is still a deeper problem: what it leaves out.
Fluid intelligence and abstract reasoning. Cattell’s distinction between crystallised intelligence (acquired knowledge) and fluid intelligence (the capacity to reason through novel problems) maps almost exactly onto this debate. A job-specific task samples crystallised, domain-bound knowledge. A graduate programme is, above all, a bet on fluid ability: the capacity to walk into a situation nobody has trained you for and work it out. General cognitive ability has been the most robust predictor of job performance across occupations for a century of research, and its predictive value is highest precisely for complex roles and for people who are still learning. A skills-based task does not measure it. It cannot, because the task is designed to look like a job rather than to isolate a psychological construct.
Personality. Conscientiousness predicts performance across virtually every job family studied, from Barrick and Mount’s 1991 meta-analysis through to the 2022 Sackett re-analysis, and it predicts it independently of ability. Emotional stability predicts resilience under pressure. Agreeableness and related traits predict how someone functions in a team, a finding replicated in Bell’s 2007 meta-analysis of team composition. These are the qualities graduate employers actually say they want: reliability, coachability, teamwork, the ability to take feedback. A mock email exercise does not measure whether a person will still be reliable in month eight, or whether they can share credit, or how they respond when a manager corrects them.
Proponents of the skills-based approach are often candid about this. The case for it typically contrasts task-based assessment with personality testing, characterises personality and soft-skill measures as a weaker or riskier proposition from a validation standpoint, and sometimes asserts outright that personality tests are not designed to predict job performance. On the evidence, that last claim is simply wrong. Well-constructed measures of the Big Five have one of the largest and most replicated evidence bases in applied psychology. An assessment process can reasonably choose not to include them. It should not be sold to HR buyers on the basis that the science is against them.
Work values and motivation. Whether a graduate stays beyond the programme is largely a question of fit between what they want from work and what the role provides. No task-based assessment addresses this at all.
Put these together and the picture is stark. For a graduate, the approach measures the one thing they have not yet had the chance to acquire, and omits the three things that best predict whether they will acquire it and stay.
The face validity trap
There is a further problem that applies to skills-based tasks regardless of the population, and it is the one HR teams are least prepared for.
Face validity means a test looks relevant to the job. It is a property of appearance, not of measurement. In the Standards for Educational and Psychological Testing published by the AERA, APA and NCME, the recognised sources of validity evidence are content, response processes, internal structure, relationships with other variables, and consequences. Face validity is not among them. A test can look perfectly job-relevant and predict nothing, and a test can look abstract and predict a great deal.
The clearest demonstration is the one the skills-based approach defines itself against. A test of abstract reasoning, with its matrices, sequences and number series, looks nothing like any job on earth. No accountant, engineer or customer service officer spends their day completing patterns. Yet general cognitive ability is the single most consistently replicated predictor of job performance in the history of selection research, across essentially every occupation studied, and it is most predictive precisely where the work is complex and the person is still learning. If face validity were evidence, that finding would be impossible. The research says the opposite: how much a test resembles the job tells you nothing about how well it predicts the job.
A claim that recurs in the case for skills-based assessment is that a test which directly mimics what a person will do on the job can, for that reason alone, be regarded as validated. It is worth pausing on that idea, because it is the clearest expression of the confusion at the heart of the approach. Mimicking the job is a claim about content. Validation is a claim about prediction. They are not the same thing, and the gap between them is where poor hiring decisions live.
Here is how the gap opens up in practice. Any job is made up of dozens of task components, and their contribution to overall performance varies enormously. Campbell’s model of job performance identifies several distinct dimensions, of which job-specific task proficiency is only one, sitting alongside effort, discipline, teamwork and the facilitation of others’ performance. A skills-based assessment samples one narrow slice of the first dimension. Its tasks are chosen because they are easy to simulate on a screen and easy to auto-grade, not because a job analysis identified them as the components that separate strong performers from weak ones.
So an assessment can be built around a task that is genuinely part of the role and still be largely irrelevant to performance in the role. A graduate consultant does write client emails. Whether they write a slightly better one under exam conditions is a poor proxy for whether they will become a good consultant. Psychometricians call this criterion deficiency: the measure captures a fragment of the performance domain and presents it as the whole. Because the fragment is so visibly job-like, nobody notices what has been left out. That is the danger of face validity. It does not just fail to add information. It actively persuades the decision-maker that no further information is needed.
A hypothetical case makes the point. Two graduates apply for the same analyst programme. The first spent a summer in a corporate finance team, is fluent in the spreadsheet conventions the assessment uses, and scores in the top decile on the task. The second worked in hospitality through university, has never built a financial model, and scores in the middle of the pack. On a task-based ranked shortlist, the first is a clear hire and the second is marginal.
Now add what a psychometric battery would have shown. The second graduate has abstract reasoning in the top few per cent, high conscientiousness and strong emotional stability. The first is average on all three. Eighteen months later, both have completed the same onboarding and both can build a model. The exposure advantage that produced the task score has been fully absorbed by training and no longer separates them. What now separates them is everything the task never measured: how fast each picks up unfamiliar work, how reliably each delivers under pressure, how each handles correction, how each functions in a team. On the evidence, the second graduate is the stronger long-term performer, and the process that looked most job-relevant is the process that would have screened them out.
That is what a job-like task costs you. It rewards a temporary, trainable advantage as though it were a durable one, and it does so most confidently in exactly the population where the advantage is least meaningful.
Has skills-based assessment been validated?
The honest answer is that, in most cases, nobody outside the organisation that built the assessment knows, and that ought to concern anyone using it to make decisions about people.
Search the literature and you will not find an independent, peer-reviewed criterion validity study of auto-generated, AI-scored skills assessments of the kind now being promoted for graduate hiring. What is typically published in support of the approach is a description of a development process, a third-party audit of the scoring algorithms for bias, statements that the methodology is grounded in industrial-organisational psychology, and claims that scoring models are trained on human-graded responses and calibrated against hiring outcomes. Every one of those is an in-house assertion. None has been subjected to the kind of external scrutiny that a technical manual, a journal article or an independent replication provides.
Before you adopt one, put the following questions to whoever is proposing it.
- Where is the peer-reviewed research showing this approach is as good as, or better than, well-researched psychometric assessment? That is the claim on which skills-first hiring rests, and it is a testable one. The test is simple: a published, independent study in which this kind of assessment was run alongside validated measures of cognitive ability and personality on the same applicants, with results compared against actual job performance. Evidence for work samples in general does not count; it has to be this assessment, scored this way. If no such study exists, the approach is being recommended over better-evidenced alternatives on the strength of assertion alone.
- What population was it validated on? If the answer is experienced hires, the evidence does not transfer to graduates. If the answer is graduates, the study should be easy to produce.
- What does “grounded in I/O psychology principles” mean in practice? Was a job analysis conducted by a qualified person for each role, or was the task list inferred by a language model from a job advertisement? These are not equivalent, and the difference is the difference between content validity and a guess.
- What is the scoring model actually measuring? If it was trained on human graders’ judgements, it has learned their preferences and their blind spots and will reproduce them with perfect consistency. Consistency is a reliability property. It is not accuracy, and it is not validity. If the model was calibrated against hire decisions rather than performance outcomes, it has learned to predict which candidates recruiters liked, which is a different thing again.
- Is a bias audit being presented as validation? An audit confirming that scores do not differ by protected group is worth having, and it is relevant to an employer’s obligations under Australian anti-discrimination law and the equivalent legislation in New Zealand. But a test can be perfectly fair and perfectly useless. Absence of adverse impact is not evidence of predictive validity, and buyers should not accept one as a substitute for the other.
- Is a registered psychologist accountable for the interpretation? In Australia, psychological testing that informs decisions about individuals is a regulated professional activity with professional standards attached. A software dashboard is not a professional, and an auto-generated ranking is not an interpretation.
Complementary, not a replacement
None of this means skills-based tasks should be thrown out. Skills testing is a complement to psychometric assessment, not a substitute for it, and the reason is exactly the one set out above: it measures declarative and procedural knowledge, which is useful, while leaving untouched the broader and better-researched predictors of how a person will learn, adapt and behave.
For graduate selection specifically, the ordering matters. Skills-based assessment should be secondary to well-researched psychometric measures, not the other way round, and the logic follows directly from everything above. The psychometric measures assess the durable, general characteristics that determine how a graduate will acquire the job: their reasoning capacity, their conscientiousness, their stability under pressure, their fit with the work. The skills task assesses a narrow, trainable, experience-dependent proficiency that the job itself will shortly provide. Putting the narrow and trainable measure first, and using it to decide who is even seen, means you filter on the thing that matters least before you have measured the things that matter most. The evidence base runs the same way: decades of replicated, peer-reviewed research behind the psychometric measures, and in-house assertions behind the auto-generated task. When one predictor is durable and well evidenced and the other is temporary and unproven, the well-evidenced one should carry the weight.
Used properly, that looks like this. A validated measure of general cognitive ability, including abstract and fluid reasoning, to establish learning capacity. A well-normed personality inventory built on the Big Five to establish conscientiousness, emotional stability and the interpersonal traits that govern teamwork. A work values or motivation measure to establish fit and likely retention. A structured interview. And then, for experienced applicants where it adds information, or late in a graduate process as a realistic job preview or a tie-breaker between candidates who have already cleared the psychometric bar, a well-designed skills task with a human professional accountable for its interpretation.
What it does not look like is your graduate programme replacing its psychometric screen with a job-description-to-assessment generator, on the strength of a sales deck, and then wondering three years later why the intake skews towards the well-connected and retention has fallen.
The cost of getting this wrong
Consider what actually happens when your graduate programme screens on job-specific task performance. The shortlist fills with candidates who had the most relevant prior exposure. That group is not random. It over-represents students who could afford unpaid internships, who attended universities with the strongest industry links, and whose networks opened doors early. A bias audit may report, quite correctly, that scores do not differ by gender or ethnicity. And the process will nonetheless have selected for socio-economic advantage while discarding high-ability candidates who simply had not seen the task before.
Meanwhile you lose the very things you set out to find: reasoning under novelty, conscientiousness, coachability, and the values alignment that determines whether a graduate stays and grows. None of these appear on a spreadsheet exercise.
The instruments with the strongest evidence for inexperienced populations remain the ones that have been studied for the longest. They are less glamorous than a simulation. They also work, and they can prove it. Fashionable is not the same as valid, and you owe it to your candidates, and to your organisation, to know the difference.
Frequently asked questions
What is skills-based assessment?
Skills-based (or skills-first) assessment presents candidates with job-style tasks, such as writing, coding or spreadsheet exercises, and scores their responses, increasingly with artificial intelligence. It is a form of work sample testing.
Why are skills-based assessments a poor fit for graduate recruitment?
Graduates have not yet had the opportunity to acquire the job-specific declarative and procedural knowledge these tasks measure, because that knowledge is normally provided through on-the-job training. Scores therefore reflect prior exposure rather than potential. The research on work samples has, since Schmidt and Hunter’s 1998 review, been explicit that they are appropriate only for applicants who already know the job.
What do skills-based tasks fail to measure?
Fluid intelligence and abstract reasoning (the capacity to learn and reason through novel problems), personality traits such as conscientiousness, emotional stability and the interpersonal traits that predict teamwork, and work values that predict retention. These are the best-supported predictors of performance for inexperienced hires.
What is the difference between face validity and criterion validity?
Face validity means a test looks relevant to the job. Criterion validity means test scores have been shown statistically to predict performance in the job. Only the second is evidence that a test works. A task can be part of a role and still be largely irrelevant to overall performance in it.
Is there published validation research for skills-based assessment?
What is typically published in support of the approach is a description of the development process, third-party bias audits of scoring algorithms, and statements that the methods are grounded in industrial-organisational psychology. As at the date of writing, no independent, peer-reviewed criterion validity study of auto-generated, AI-scored skills assessments appears to have been published. A bias audit is a fairness check, not evidence of predictive validity.
Should skills-based tasks be used at all?
Yes, as a complement to validated psychometric measures: with experienced applicants, or late in a graduate process as a realistic job preview, with a registered psychologist accountable for interpretation.
Should skills-based assessment be primary or secondary in graduate selection?
Secondary. Well-researched psychometric measures of cognitive ability, personality and work values assess the durable characteristics that determine how a graduate will learn the job, and they carry decades of peer-reviewed evidence. A skills-based task assesses a narrow, trainable proficiency that the job will soon provide, and rests on in-house evidence. The well-evidenced, durable predictors should decide who progresses; the skills task can then add information at a later stage.
Which assessments are best for graduate programmes?
Validated measures of general cognitive ability including abstract reasoning, a well-normed Big Five personality inventory, a work values or motivation measure, and a structured interview provide the strongest evidence base for predicting graduate performance and retention.
References
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for Educational and Psychological Testing. AERA.
- Anderson, J. R. (1982). Acquisition of cognitive skill. Psychological Review, 89(4), 369–406.
- Barrick, M. R., & Mount, M. K. (1991). The Big Five personality dimensions and job performance: A meta-analysis. Personnel Psychology, 44(1), 1–26.
- Bell, S. T. (2007). Deep-level composition variables as predictors of team performance: A meta-analysis. Journal of Applied Psychology, 92(3), 595–615.
- Campbell, J. P., McCloy, R. A., Oppler, S. H., & Sager, C. E. (1993). A theory of performance. In N. Schmitt & W. C. Borman (Eds.), Personnel selection in organizations (pp. 35–70). Jossey-Bass.
- Cattell, R. B. (1963). Theory of fluid and crystallized intelligence: A critical experiment. Journal of Educational Psychology, 54(1), 1–22.
- Kanfer, R., & Ackerman, P. L. (1989). Motivation and cognitive abilities: An integrative/aptitude-treatment interaction approach to skill acquisition. Journal of Applied Psychology, 74(4), 657–690.
- Roth, P. L., Bobko, P., & McFarland, L. A. (2005). A meta-analysis of work sample test validity: Updating and integrating some classic literature. Personnel Psychology, 58(4), 1009–1037.
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
- Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274.
- Schmidt, F. L., Hunter, J. E., & Outerbridge, A. N. (1986). Impact of job experience and ability on job knowledge, work sample performance, and supervisory ratings of job performance. Journal of Applied Psychology, 71(3), 432–439.