Every now and again an older idea circles back around in selection psychology and finds a new audience. The current example is skills-based assessment, sometimes called “skills-first” hiring: the idea that the best way to select people is to have them complete job-style tasks, typically auto-generated from a job description and scored by artificial intelligence, and to rank them on the result. It is increasingly promoted for early-careers hiring.
The novelty is in the packaging. Strip away the interface and the algorithm and what remains is the work sample test, a method selection psychologists have been building, studying and refining for the better part of a century. The tasks are now generated automatically rather than designed by a job analyst, scored by a model rather than a trained assessor, and marketed as a break from the psychometric tradition rather than what it actually is: a well-worn branch of it, revived with the caveats left off.
To be clear about what this article is not arguing: assessing what people can do, rather than what their CV says, is a sound idea. That is precisely why work sample testing has survived for so long, and skills testing has a legitimate place in a well-designed process. The concern here is narrower and more serious. The approach is now being promoted specifically for graduate and early-careers recruitment, and that is the one population for which its entire logic breaks down. If you run a graduate programme, or advise one, this matters to you.
The graduate paradox
A work sample test rests on a simple premise: if you want to know whether someone can do the job, watch them do a piece of it. That premise holds when the applicant has done the job, or something very like it, before. It collapses when they have not.
A graduate applying for their first analyst, consulting or customer-facing role has, by definition, not yet done the work. The spreadsheet exercise, the mock stakeholder email, the ticket-triage simulation, the short business case: these are precisely the tasks the organisation will teach them in their first weeks of onboarding. Look at your own programme’s induction and you will almost certainly find it built around exactly this material. So the assessment asks candidates to demonstrate, before they are hired, the skills you are about to provide as part of the job.
What does a score on such a task actually tell you about a twenty-two-year-old? It tells you who has already been exposed to that specific kind of work: the student who landed the right internship, whose university ran the right software, whose family happened to be in the industry. It does not tell you who will be the strongest performer once everyone has been through the same training. Those are different questions, and for graduates only the second one matters.
Advocates of skills-based assessment commonly present role-specific tasks as well suited to graduates because they lack prior work experience: the argument runs that a task levels the field where a CV cannot. The opposite is true. The absence of prior experience is exactly the condition under which a work sample stops measuring capability and starts measuring exposure.
What the research says about experience and work samples
This is not a novel objection. It has been stated in the literature since the modern era of selection research began.
Schmidt and Hunter’s 1998 review, still the most cited paper in the field, was explicit that job knowledge tests and work samples are only appropriate for applicants who already know the job, and are of no use with inexperienced applicants. Roth, Bobko and McFarland, in their 2005 meta-analysis of work sample validity, made the same caveat: the good validity figures come from studies of experienced applicants. In the 2022 re-analysis by Sackett, Zhang, Berry and Lievens, which corrected the older literature for statistical over-adjustment, work samples and job knowledge tests remained respectable predictors, but nothing in that paper extends the finding to people who have never done the work.
Why does experience matter so much? Schmidt, Hunter and Outerbridge showed in 1986 that job performance is driven by job knowledge, and that job knowledge is driven by two things: general cognitive ability and time on the job. Ability determines how fast and how completely a person learns the work. Experience provides the opportunity. A job knowledge or work sample test measures the output of that learning process. For an experienced applicant the output is a fair summary of what they have acquired. For a graduate the output does not yet exist, so the sensible thing to measure is the input: the ability, and the dispositions, that determine how well they will learn.
Cognitive psychology reaches the same conclusion from a different direction. John Anderson’s work on skill acquisition distinguishes declarative knowledge (knowing that: facts, rules, procedures you can describe) from procedural knowledge (knowing how: skills that have become fluent through practice). Kanfer and Ackerman showed that early in learning any new skill, performance depends heavily on general cognitive ability, because the learner is still building the declarative foundation; only after extended practice does performance become proceduralised and depend on the specific experience gained. Skills-based tasks tap declarative and procedural knowledge of a specific job. Graduates are, by definition, at the start of that curve. The construct the approach measures is the one construct graduates cannot yet be expected to have.
What skills-based assessment cannot see
Suppose the task is well built and fairly scored. There is still a deeper problem: what it leaves out.
Fluid intelligence and abstract reasoning. Cattell’s distinction between crystallised intelligence (acquired knowledge) and fluid intelligence (the capacity to reason through novel problems) maps almost exactly onto this debate. A job-specific task samples crystallised, domain-bound knowledge. A graduate programme is, above all, a bet on fluid ability: the capacity to walk into a situation nobody has trained you for and work it out. General cognitive ability has been the most robust predictor of job performance across occupations for a century of research, and its predictive value is highest precisely for complex roles and for people who are still learning. A skills-based task does not measure it. It cannot, because the task is designed to look like a job rather than to isolate a psychological construct.
Personality. Conscientiousness predicts performance across virtually every job family studied, from Barrick and Mount’s 1991 meta-analysis through to the 2022 Sackett re-analysis, and it predicts it independently of ability. Emotional stability predicts resilience under pressure. Agreeableness and related traits predict how someone functions in a team, a finding replicated in Bell’s 2007 meta-analysis of team composition. These are the qualities graduate employers actually say they want: reliability, coachability, teamwork, the ability to take feedback. A mock email exercise does not measure whether a person will still be reliable in month eight, or whether they can share credit, or how they respond when a manager corrects them.
Proponents of the skills-based approach are often candid about this. The case for it typically contrasts task-based assessment with personality testing, characterises personality and soft-skill measures as a weaker or riskier proposition from a validation standpoint, and sometimes asserts outright that personality tests are not designed to predict job performance. On the evidence, that last claim is simply wrong. Well-constructed measures of the Big Five have one of the largest and most replicated evidence bases in applied psychology. An assessment process can reasonably choose not to include them. It should not be sold to HR buyers on the basis that the science is against them.
Work values and motivation. Whether a graduate stays beyond the programme is largely a question of fit between what they want from work and what the role provides. No task-based assessment addresses this at all.
Put these together and the picture is stark. For a graduate, the approach measures the one thing they have not yet had the chance to acquire, and omits the three things that best predict whether they will acquire it and stay.
The face validity trap
There is a further problem that applies to skills-based tasks regardless of the population, and it is the one HR teams are least prepared for.
Face validity means a test looks relevant to the job. It is a property of appearance, not of measurement. In the Standards for Educational and Psychological Testing published by the AERA, APA and NCME, the recognised sources of validity evidence are content, response processes, internal structure, relationships with other variables, and consequences. Face validity is not among them. A test can look perfectly job-relevant and predict nothing, and a test can look abstract and predict a great deal.
The clearest demonstration is the one the skills-based approach defines itself against. A test of abstract reasoning, with its matrices, sequences and number series, looks nothing like any job on earth. No accountant, engineer or customer service officer spends their day completing patterns. Yet general cognitive ability is the single most consistently replicated predictor of job performance in the history of selection research, across essentially every occupation studied, and it is most predictive precisely where the work is complex and the person is still learning. If face validity were evidence, that finding would be impossible. The research says the opposite: how much a test resembles the job tells you nothing about how well it predicts the job.
A claim that recurs in the case for skills-based assessment is that a test which directly mimics what a person will do on the job can, for that reason alone, be regarded as validated. It is worth pausing on that idea, because it is the clearest expression of the confusion at the heart of the approach. Mimicking the job is a claim about content. Validation is a claim about prediction. They are not the same thing, and the gap between them is where poor hiring decisions live.
Here is how the gap opens up in practice. Any job is made up of dozens of task components, and their contribution to overall performance varies enormously. Campbell’s model of job performance identifies several distinct dimensions, of which job-specific task proficiency is only one, sitting alongside effort, discipline, teamwork and the facilitation of others’ performance. A skills-based assessment samples one narrow slice of the first dimension. Its tasks are chosen because they are easy to simulate on a screen and easy to auto-grade, not because a job analysis identified them as the components that separate strong performers from weak ones.
So an assessment can be built around a task that is genuinely part of the role and still be largely irrelevant to performance in the role. A graduate consultant does write client emails. Whether they write a slightly better one under exam conditions is a poor proxy for whether they will become a good consultant. Psychometricians call this criterion deficiency: the measure captures a fragment of the performance domain and presents it as the whole. Because the fragment is so visibly job-like, nobody notices what has been left out. That is the danger of face validity. It does not just fail to add information. It actively persuades the decision-maker that no further information is needed.
A hypothetical case makes the point. Two graduates apply for the same analyst programme. The first spent a summer in a corporate finance team, is fluent in the spreadsheet conventions the assessment uses, and scores in the top decile on the task. The second worked in hospitality through university, has never built a financial model, and scores in the middle of the pack. On a task-based ranked shortlist, the first is a clear hire and the second is marginal.
Now add what a psychometric battery would have shown. The second graduate has abstract reasoning in the top few per cent, high conscientiousness and strong emotional stability. The first is average on all three. Eighteen months later, both have completed the same onboarding and both can build a model. The exposure advantage that produced the task score has been fully absorbed by training and no longer separates them. What now separates them is everything the task never measured: how fast each picks up unfamiliar work, how reliably each delivers under pressure, how each handles correction, how each functions in a team. On the evidence, the second graduate is the stronger long-term performer, and the process that looked most job-relevant is the process that would have screened them out.
That is what a job-like task costs you. It rewards a temporary, trainable advantage as though it were a durable one, and it does so most confidently in exactly the population where the advantage is least meaningful.
The reliability problem behind the validity claims
Before any test can be shown to be valid, it has to be shown to be reliable. Reliability is not a footnote to validity; it is a precondition for it. On classical psychometric principles, a test’s correlation with an outside criterion, its validity coefficient, cannot exceed the square root of its own reliability coefficient. An instrument with unknown reliability does not have an unproven validity claim. It has an untestable one.
Reliability is established in a small number of recognised ways, and every one of them depends on the same thing: a fixed, stable item or item set, answered by a sufficiently large sample of people, so that a statistic can be computed. Internal consistency, most commonly estimated with Cronbach’s alpha, requires the same items answered by many candidates so that item-to-item consistency can be measured; the older equivalent for right-or-wrong scoring is KR-20, and the original method, split-half reliability, works the same way by correlating one half of a fixed test against the other. Test-retest reliability requires the same instrument given to the same people on two occasions, with scores correlated to check stability over time. Parallel-forms reliability requires two equivalent versions of a test, given to the same people, correlated against each other. Each method is mature and well understood. Each requires an item to exist, in the same form, for long enough to be measured against itself or against many other people’s responses to it.
This is where an assessment generated fresh for every job description runs into a structural problem, not just an evidential gap. If each set of tasks is assembled on the spot, from a new job description, rather than drawn from a stable, calibrated bank, then no single item has been answered by a large sample of people in that exact form. There is nothing to split in half, nothing to retest, no parallel form to compare it with. Reliability, on any of the classical methods, cannot be computed for an instrument that does not stay the same for long enough to be measured against itself.
There is also a practical threshold that sits alongside this theoretical one. A widely used rule of thumb in test development holds that an item needs to have been completed by at least around 100 candidates before reliability statistics such as Cronbach’s alpha become stable enough to interpret, and many methodologists prefer several hundred for real precision. An item written uniquely for a single job description, and reassembled or reworded for the next, will rarely reach that number of prior respondents in its exact form. Statistically speaking, it is untested at the moment it is first shown to a candidate, and it may remain untested indefinitely if it is never repeated.
This is not a criticism of how well the underlying model is built. A grading model can be trained on a large volume of human-graded responses, and that is a genuine claim about the model’s general behaviour. It is not the same as a claim about the reliability of any one generated assessment, because reliability is a property of a specific instrument given to a specific population, not a property of the software that produced it.
There is one honest exception, and it is worth stating so as not to overclaim. If the tasks are in fact drawn from a pre-calibrated item bank, where each item’s properties have already been established using item response theory across a large sample, reliability has a genuine technical basis, and the process is closer to adaptive testing than to true on-the-fly item writing. Ask which of these two things is actually happening before assuming the worse case. Marketing language that emphasises assessments “generated in seconds” with “no fixed question bank” points towards the first, unmeasurable version, not the second.
The consequence for validity is direct, and it is the point this whole article turns on. A validity claim about an instrument whose reliability cannot be established is not simply unproven. It is unfalsifiable: there is no coherent way to test it. That is a harder problem than a missing study, because a missing study could in principle be commissioned tomorrow. A structural inability to compute reliability is a property of the design itself, and no amount of additional marketing closes that gap on its own.
This is also where the risk to buyers concentrates most sharply. Words like “validated”, “AI-powered” and “grounded in I/O psychology” read as reassuring to a hiring manager without psychometric training, precisely because they sound scientific. Face validity, discussed above, is one part of that impression. An unstated and unmeasured reliability is the other, and it is the harder one to notice, because nobody outside a testing background thinks to ask for it.
Has skills-based assessment been validated?
The honest answer is that, in most cases, nobody outside the organisation that built the assessment knows, and that ought to concern anyone using it to make decisions about people.
Search the literature and you will not find an independent, peer-reviewed criterion validity study of auto-generated, AI-scored skills assessments of the kind now being promoted for graduate hiring. What is typically published in support of the approach is a description of a development process, a third-party audit of the scoring algorithms for bias, statements that the methodology is grounded in industrial-organisational psychology, and claims that scoring models are trained on human-graded responses and calibrated against hiring outcomes. Every one of those is an in-house assertion. None has been subjected to the kind of external scrutiny that a technical manual, a journal article or an independent replication provides.
Before you adopt one, put the following questions to whoever is proposing it.
- Has reliability been established for this specific assessment, and how? Reliability comes before validity, not alongside it. Ask which method was used, internal consistency, test-retest or otherwise, and on how many candidates. As a rule of thumb, an item needs at least around 100 completions before reliability statistics are considered stable enough to interpret. If tasks are generated fresh for every job description with no stable item bank behind them, ask whether any single item has ever reached that number, let alone whether reliability was formally computed.
- Where is the peer-reviewed research showing this approach is as good as, or better than, well-researched psychometric assessment? That is the claim on which skills-first hiring rests, and it is a testable one. The test is simple: a published, independent study in which this kind of assessment was run alongside validated measures of cognitive ability and personality on the same applicants, with results compared against actual job performance. Evidence for work samples in general does not count; it has to be this assessment, scored this way. If no such study exists, the approach is being recommended over better-evidenced alternatives on the strength of assertion alone.
- What population was it validated on? If the answer is experienced hires, the evidence does not transfer to graduates. If the answer is graduates, the study should be easy to produce.
- What does “grounded in I/O psychology principles” mean in practice? Was a job analysis conducted by a qualified person for each role, or was the task list inferred by a language model from a job advertisement? These are not equivalent, and the difference is the difference between content validity and a guess.
- What is the scoring model actually measuring? If it was trained on human graders’ judgements, it has learned their preferences and their blind spots and will reproduce them with perfect consistency. Consistency is a reliability property. It is not accuracy, and it is not validity. If the model was calibrated against hire decisions rather than performance outcomes, it has learned to predict which candidates recruiters liked, which is a different thing again.
- Is a bias audit being presented as validation? An audit confirming that scores do not differ by protected group is worth having, and it is relevant to an employer’s obligations under Australian anti-discrimination law and the equivalent legislation in New Zealand. But a test can be perfectly fair and perfectly useless. Absence of adverse impact is not evidence of predictive validity, and buyers should not accept one as a substitute for the other.
- Is a registered psychologist accountable for the interpretation? In Australia, psychological testing that informs decisions about individuals is a regulated professional activity with professional standards attached. A software dashboard is not a professional, and an auto-generated ranking is not an interpretation.
What, exactly, is being purchased?
There is a question worth asking before any contract is signed, and it has nothing to do with reliability or validity: what capability is actually being bought?
Turning a job description into a set of job-style tasks, and asking a general-purpose AI model to draft and mark them, is not a specialised skill held by a small number of vendors. Any HR team with access to a mainstream AI assistant can paste in a real position description today and ask it to propose written tasks, spreadsheet exercises, case scenarios or interview-style questions built around that role. This is, in substance, the starting point these platforms are built on. What a vendor typically adds on top is delivery infrastructure: a candidate-facing interface, a scoring workflow and a reporting dashboard. That is software packaging, not psychometric science, and none of it makes the resulting scores any more reliable or any more valid.
This distinction matters because it clarifies what is, and is not, being paid for. An organisation paying a premium for a “validated” skills-based assessment may, at the level of the actual measurement content, be paying for the packaging around a capability it already has access to, described in language borrowed from a discipline the product itself has never been tested against. If the task-generation step can be reproduced by any manager with a general-purpose AI assistant and ten minutes to spare, the case for treating the output as a scientifically credentialed instrument gets weaker, not stronger.
This is not an argument for building a do-it-yourself version instead. The reliability and validity problems set out above apply equally, whichever party generates the tasks. It is an argument for pricing and evaluating the product accurately. Before signing anything, it is worth testing this directly: take a real job description for the role in question and ask a mainstream AI assistant to draft three or four job-style tasks for it, then compare the result with what the vendor’s demonstration produces. If the two look substantially alike, the premium being paid is for hosting, workflow and a dashboard, not for a scientifically distinct measurement method, and it should be evaluated on those terms.
Complementary, not a replacement
None of this means skills-based tasks should be thrown out. Skills testing is a complement to psychometric assessment, not a substitute for it, and the reason is exactly the one set out above: it measures declarative and procedural knowledge, which is useful, while leaving untouched the broader and better-researched predictors of how a person will learn, adapt and behave.
For graduate selection specifically, the ordering matters. Skills-based assessment should be secondary to well-researched psychometric measures, not the other way round, and the logic follows directly from everything above. The psychometric measures assess the durable, general characteristics that determine how a graduate will acquire the job: their reasoning capacity, their conscientiousness, their stability under pressure, their fit with the work. The skills task assesses a narrow, trainable, experience-dependent proficiency that the job itself will shortly provide. Putting the narrow and trainable measure first, and using it to decide who is even seen, means you filter on the thing that matters least before you have measured the things that matter most. The evidence base runs the same way: decades of replicated, peer-reviewed research behind the psychometric measures, and in-house assertions behind the auto-generated task. When one predictor is durable and well evidenced and the other is temporary and unproven, the well-evidenced one should carry the weight.
Used properly, that looks like this. A validated measure of general cognitive ability, including abstract and fluid reasoning, to establish learning capacity. A well-normed personality inventory built on the Big Five to establish conscientiousness, emotional stability and the interpersonal traits that govern teamwork. A work values or motivation measure to establish fit and likely retention. A structured interview. And then, for experienced applicants where it adds information, or late in a graduate process as a realistic job preview or a tie-breaker between candidates who have already cleared the psychometric bar, a well-designed skills task with a human professional accountable for its interpretation.
What it does not look like is your graduate programme replacing its psychometric screen with a job-description-to-assessment generator, on the strength of a sales deck, and then wondering three years later why the intake skews towards the well-connected and retention has fallen.
The exposure this creates in government recruitment
This is not only a scientific problem. In Australian government recruitment, merit is not a preference; it is a legislated requirement, and that requirement gives a rejected candidate, or an independent reviewer, a real avenue to challenge a decision built on an unvalidated, unreliable instrument.
Section 10A of the Public Service Act 1999 requires Australian Public Service recruitment decisions to be based on merit against the genuine requirements of the role. The Merit Protection Commissioner’s own guidance to agencies is direct on this point: agencies must be able to demonstrate the effectiveness of any AI-assisted tool in assessing candidates against criteria that reflect the work-related qualities genuinely required for the role. An agency that cannot answer the questions above, no reliability evidence, no validity evidence, no documented job analysis behind the task list, has no answer to that requirement either.
This is not a hypothetical risk. In its 2021–22 annual report, the Merit Protection Commissioner reviewed a bulk AI-assisted recruitment round and found the process had failed to select the right people for the job; that year the Commissioner overturned twelve recruitment decisions, eleven from that single round. Union representatives told a parliamentary committee in December 2024 that comparable tools had rated experienced Services Australia staff unsuitable for roles and promotions because their applications had not used the language an algorithm was searching for. In April 2026 the Australian Public Service Commission issued new principles and guidance on AI in recruitment, stating plainly that AI use must not replace human judgement, must not compromise merit-based selection, and that agencies remain fully accountable and cannot delegate the decision. The guidance names review by the Merit Protection Commissioner, and discrimination complaints arising from algorithmic bias, as the specific consequences of getting this wrong. State jurisdictions apply the same logic: in New South Wales, Rule 16 of the Government Sector Employment (General) Rules 2014 sets out the merit principle governing every public service employment decision, and the state’s Public Service Commission has issued its own guidance addressed specifically at AI recruitment tools. New Zealand’s public service operates under an equivalent good-employer and merit-based appointment framework, with its own avenues for review.
None of this requires a candidate to be litigious. A capable, discerning graduate applicant, increasingly aware of how these tools work and increasingly willing to ask, can simply request the basis on which they were screened out. In the public sector, that request has a formal channel to travel through. An employer that cannot produce reliability evidence, validity evidence or a documented job analysis behind the task list is not well placed to defend the decision, whichever direction the challenge comes from, and the exposure sits with the agency, not with the technology it purchased.
This is general information, not legal advice, and organisations should seek their own advice on the requirements that apply to their recruitment processes. The practical point holds regardless of jurisdiction: an assessment method that cannot answer the questions above is a method that cannot be defended when someone asks.
The cost of getting this wrong
The legal exposure above is one cost, and it falls on the organisation. There is a second cost, and it falls on the candidates the process is supposed to serve. Consider what actually happens when your graduate programme screens on job-specific task performance. The shortlist fills with candidates who had the most relevant prior exposure. That group is not random. It over-represents students who could afford unpaid internships, who attended universities with the strongest industry links, and whose networks opened doors early. A bias audit may report, quite correctly, that scores do not differ by gender or ethnicity. And the process will nonetheless have selected for socio-economic advantage while discarding high-ability candidates who simply had not seen the task before.
Meanwhile you lose the very things you set out to find: reasoning under novelty, conscientiousness, coachability, and the values alignment that determines whether a graduate stays and grows. None of these appear on a spreadsheet exercise.
The instruments with the strongest evidence for inexperienced populations remain the ones that have been studied for the longest. They are less glamorous than a simulation. They also work, and they can prove it. Fashionable is not the same as valid, and you owe it to your candidates, and to your organisation, to know the difference.
Frequently asked questions
What is skills-based assessment?
Skills-based (or skills-first) assessment presents candidates with job-style tasks, such as writing, coding or spreadsheet exercises, and scores their responses, increasingly with artificial intelligence. It is a form of work sample testing.
Why are skills-based assessments a poor fit for graduate recruitment?
Graduates have not yet had the opportunity to acquire the job-specific declarative and procedural knowledge these tasks measure, because that knowledge is normally provided through on-the-job training. Scores therefore reflect prior exposure rather than potential. The research on work samples has, since Schmidt and Hunter’s 1998 review, been explicit that they are appropriate only for applicants who already know the job.
What do skills-based tasks fail to measure?
Fluid intelligence and abstract reasoning (the capacity to learn and reason through novel problems), personality traits such as conscientiousness, emotional stability and the interpersonal traits that predict teamwork, and work values that predict retention. These are the best-supported predictors of performance for inexperienced hires.
What is the difference between face validity and criterion validity?
Face validity means a test looks relevant to the job. Criterion validity means test scores have been shown statistically to predict performance in the job. Only the second is evidence that a test works. A task can be part of a role and still be largely irrelevant to overall performance in it.
Can an assessment that generates new tasks for every job description be reliable?
Reliability is a precondition for validity: an instrument’s correlation with job performance cannot exceed the square root of its own reliability. Classical reliability methods, such as internal consistency (Cronbach’s alpha) and test-retest correlation, all require a fixed item or item set answered by a large sample of candidates, with a widely used rule of thumb setting the minimum at around 100 completions per item before the statistics are considered stable. If tasks are generated fresh for each role with no stable item bank behind them, individual items rarely if ever reach that number of prior respondents, so there is no accepted way to compute reliability for that specific assessment, and any validity claim resting on it cannot be properly tested either.
Is there published validation research for skills-based assessment?
What is typically published in support of the approach is a description of the development process, third-party bias audits of scoring algorithms, and statements that the methods are grounded in industrial-organisational psychology. As at the date of writing, no independent, peer-reviewed criterion validity study of auto-generated, AI-scored skills assessments appears to have been published. A bias audit is a fairness check, not evidence of predictive validity.
Can a skills-based assessment decision be challenged in government recruitment?
Potentially, yes. Australian Public Service recruitment must be based on merit under section 10A of the Public Service Act 1999, and agencies must be able to show that any AI-assisted tool assesses candidates against criteria genuinely required for the role. The Merit Protection Commissioner reviews recruitment decisions and has previously overturned a bulk round after finding an AI-assisted process failed to select the right people; the Australian Public Service Commission issued fresh guidance on this exact risk in April 2026. Equivalent merit principles apply in state public services and in New Zealand’s public service. This is general information, not legal advice; organisations should seek their own advice on their specific obligations.
Can an organisation produce the same kind of skills-based tasks without buying a platform?
For the item-generation step, in most cases, yes. Asking a general-purpose AI assistant to draft job-style tasks from a position description is not a specialised or proprietary capability. What a vendor typically adds is delivery infrastructure — a candidate interface, anti-cheating controls, a scoring workflow and reporting. That is software packaging, not evidence of psychometric expertise or of a scientifically superior measurement method.
Should skills-based tasks be used at all?
Yes, as a complement to validated psychometric measures: with experienced applicants, or late in a graduate process as a realistic job preview, with a registered psychologist accountable for interpretation.
Should skills-based assessment be primary or secondary in graduate selection?
Secondary. Well-researched psychometric measures of cognitive ability, personality and work values assess the durable characteristics that determine how a graduate will learn the job, and they carry decades of peer-reviewed evidence. A skills-based task assesses a narrow, trainable proficiency that the job will soon provide, and rests on in-house evidence. The well-evidenced, durable predictors should decide who progresses; the skills task can then add information at a later stage.
Which assessments are best for graduate programmes?
Validated measures of general cognitive ability including abstract reasoning, a well-normed Big Five personality inventory, a work values or motivation measure, and a structured interview provide the strongest evidence base for predicting graduate performance and retention.
References
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for Educational and Psychological Testing. AERA.
- Anderson, J. R. (1982). Acquisition of cognitive skill. Psychological Review, 89(4), 369–406.
- Barrick, M. R., & Mount, M. K. (1991). The Big Five personality dimensions and job performance: A meta-analysis. Personnel Psychology, 44(1), 1–26.
- Bell, S. T. (2007). Deep-level composition variables as predictors of team performance: A meta-analysis. Journal of Applied Psychology, 92(3), 595–615.
- Campbell, J. P., McCloy, R. A., Oppler, S. H., & Sager, C. E. (1993). A theory of performance. In N. Schmitt & W. C. Borman (Eds.), Personnel selection in organizations (pp. 35–70). Jossey-Bass.
- Cattell, R. B. (1963). Theory of fluid and crystallized intelligence: A critical experiment. Journal of Educational Psychology, 54(1), 1–22.
- Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297–334.
- DeVellis, R. F. (2016). Scale development: Theory and applications (4th ed.). SAGE Publications.
- Government Sector Employment (General) Rules 2014 (NSW) r 16.
- Kanfer, R., & Ackerman, P. L. (1989). Motivation and cognitive abilities: An integrative/aptitude-treatment interaction approach to skill acquisition. Journal of Applied Psychology, 74(4), 657–690.
- Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). McGraw-Hill.
- Merit Protection Commissioner. (2022). Guidance: Using AI-assisted and automated technologies in public sector recruitment. Australian Public Service Commission. https://www.mpc.gov.au/sites/default/files/2022-11/Guidance%20-%20AI-assisted%20and%20automated%20technologies%20in%20public%20sector%20recruitment.pdf
- NSW Public Service Commission. (2025). Commissioner’s Spotlight: Use of AI in recruitment. Office of the Public Service Commissioner. https://www.psc.nsw.gov.au/assets/psc/Commissioners-Spotlight-on-the-use-of-AI-in-recruitment.pdf
- Public Service Act 1999 (Cth) s 10A.
- Roth, P. L., Bobko, P., & McFarland, L. A. (2005). A meta-analysis of work sample test validity: Updating and integrating some classic literature. Personnel Psychology, 58(4), 1009–1037.
- Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
- Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274.
- Schmidt, F. L., Hunter, J. E., & Outerbridge, A. N. (1986). Impact of job experience and ability on job knowledge, work sample performance, and supervisory ratings of job performance. Journal of Applied Psychology, 71(3), 432–439.