genArete: Learner-Centered Skill Assessment
Presented by Mark Malady, BCBA
In this first talk of a three-part series, Mark Malady, BCBA, examines the state of skill-based assessment for autistic learners and other people with disabilities. He reviews recent literature showing an inverse relationship between utilization and evidence: widely used behavior-analytic instruments such as the VB-MAPP and ABLLS-R have only emerging empirical support, while less-used tools like the ABLA carry stronger validation, with funder requirements driving much of the imbalance. Drawing on CASP/APBA assessment guidelines, he unpacks key distinctions practitioners should know before selecting a tool: direct versus indirect, standardized versus non-standardized, norm-referenced versus criterion-referenced, and single-modal versus multimodal approaches, with attention to how norm-referenced comparisons can embed normalization agendas. Using the genArete assessment as a working example, he presents design features that support learner-centered practice: eliminating cold probes in favor of measurement systems that distinguish acquisition from mastery states, top-down and bottom-up entry points that respect the learner's time, varied measurement systems matched to targets, built-in assent and accommodations, and milestone-based comparison criteria benchmarked against competent performers rather than same-age norms. He closes with practical steps any clinician can apply to existing instruments, including semi-structured intake interviews to individualize assessment packages, auditing assessment selection for true differentiation across learners, separating instructional targets from ongoing environmental supports, and strategies for getting emerging instruments approved by funders.
Register free & watch nowCancel any time. Cert delivered to your inbox the moment you pass the quiz.
Learner ratings
from 52 learners
96% would recommend this CEU to other professionals in their field (23 responses)
“very good”
— Ana S.
About this CEU
In this first talk of a three-part series, Mark Malady, BCBA, examines the state of skill-based assessment for autistic learners and other people with disabilities. He reviews recent literature showing an inverse relationship between utilization and evidence: widely used behavior-analytic instruments such as the VB-MAPP and ABLLS-R have only emerging empirical support, while less-used tools like the ABLA carry stronger validation, with funder requirements driving much of the imbalance. Drawing on CASP/APBA assessment guidelines, he unpacks key distinctions practitioners should know before selecting a tool: direct versus indirect, standardized versus non-standardized, norm-referenced versus criterion-referenced, and single-modal versus multimodal approaches, with attention to how norm-referenced comparisons can embed normalization agendas. Using the genArete assessment as a working example, he presents design features that support learner-centered practice: eliminating cold probes in favor of measurement systems that distinguish acquisition from mastery states, top-down and bottom-up entry points that respect the learner's time, varied measurement systems matched to targets, built-in assent and accommodations, and milestone-based comparison criteria benchmarked against competent performers rather than same-age norms. He closes with practical steps any clinician can apply to existing instruments, including semi-structured intake interviews to individualize assessment packages, auditing assessment selection for true differentiation across learners, separating instructional targets from ongoing environmental supports, and strategies for getting emerging instruments approved by funders.
From the talk
What was covered
Why the most-used skill assessments have the thinnest evidence, and how to individualize assessment without breaking fidelity.
- Check where your main tool sits on the evidence table before your next intake: support for the VB-MAPP and ABLLS-R is still rated emerging.
- Chart the assessment package you ran for each of your last 10 learners; one color means your multimodal approach is not real yet.
- Use a semi-structured intake interview to pick which tools to run, which to skip, and write down the reason for each choice.
- genArete skips cold probes entirely and uses its measurement system to separate acquisition from mastery instead.
- Build comparison criteria from competent performers tied to the learner's actual goal instead of same-age norms.
- Ask your funder for the written policy on evidence-based assessment; the common bar is one validity study and one reliability study.
The Evidence Gap in Common Skill Assessments
Most of us pick a skill assessment (a check of what someone can do) out of habit. The funder asks for one tool. The clinic already owns it. So we run it again. This talk tested that habit against the research. A 2025 review by Banda and Hart placed the common tools in one table. Each row showed the age range, the type, the time to give it, and the research base. Two columns matter most here: reliability (do scores repeat across raters) and validity (does it measure the right thing).
The VB-MAPP landed in the emerging column. So did the ABLLS-R. Emerging means the support is real but thin. Smaller tools did better. The ABLA, from Kerr et al. in 1977, has strong support. It is a basic test of discrimination skills (telling two things apart). The EFL from McGreevy et al. also held up well. Then the picture flips. A survey of about 1,400 behavior analysts asked what people actually use. The VB-MAPP held the largest share in 2020. The ABLLS-R came next. Only four tools cleared the ten percent mark. Another 115 instruments sat below that line.
Real-world use runs opposite to the evidence. The weaker tools get used the most. One big reason is money. Funders in the American autism market pay for these tools. Some name them outright. Most are also legacy tools, over ten years old. The field is trying to close the gap. A psychometric study (how well a test measures) published this year pulled records on more than 41,000 people to check the ABLLS-R and the AFLS. Our practice ran ahead of the validation work, because clinicians needed tools right then.
We are not in alignment on our current tools for skill-based assessment being empirically validated as instruments in and of themselves.
From the talk — Mark Malady, BCBA
Four Distinctions from the CASP and APBA Assessment Guidelines
The Council of Autism Service Providers and the Association of Professional Behavior Analysts built joint guidance for assessing autistic learners. It names four splits every clinician should be able to explain out loud. The first is direct versus indirect. Direct means you watch the skill happen live. Indirect means you ask someone about it instead. You interview a parent, a teacher, or the learner.
The second split is standardized versus non-standardized (given the same way each time). The guidance defines it by how much the items can vary in delivery. It also counts how much science backs that delivery style. Malady is not the biggest fan of that definition. It is still the one the field is working from.
The third split is the one to sit with. Norm-referenced (compared to same-age peers) means the items came from a comparison group. Criterion-referenced (compared to a fixed standard) means the items came from a logical order of what mastery takes. The guidance lists the limits of each. The norm-referenced limit is blunt. The score sits next to same-age peers and ignores the person's circumstances.
The fourth split is single-modal versus multimodal (more than one tool together). The panel could not find a single instrument that fits every learner. Their advice is to build an assessment package for each person.
So a big recommendation of the guidelines is that to this point, they have not seen single instruments that can fit the need across learners.
From the talk — Mark Malady, BCBA
Where a Normalization Agenda Hides Inside an Instrument
Norm samples usually come from age groups. They rarely come from people who share the learner's diagnosis. For autistic learners that creates a real problem. Autistic people often report using different skills to reach the same outcome. A same-age norm cannot see that. Critics of our field say ABA runs on indistinguishability (looking much like non-autistic peers). Heavy use of norm-referenced tools makes that claim easy to defend.
Graphs carry the same message quietly. Picture the displays you already know: leveled bar graphs, a single bar graph, triangles that climb with age. The design tells you to fill every box by a set age. For a group as varied as autistic learners, that goal drifts straight toward sameness.
Criterion-referenced tools carry their own risk. The guidance calls it teaching to the test. Programming shrinks down to the items on the form. Build a check into quality control so that does not happen. Malady asks six questions of every tool he uses. Who is the target audience? What assumptions sit under it? What research backs it right now? How can it be misused? Where do teams hit roadblocks? Which learner profiles does it capture well?
And in each of these, the issue that I have with them is that the goal is to fill up all of the boxes.
From the talk — Mark Malady, BCBA
What a Skill Assessment Should Actually Give You
Before you judge a tool, decide what the job is. The talk listed six outcomes a skill assessment should deliver. It should give an accurate account of what the learner can do. It should not credit skills the person lacks. It should not miss skills the person has. It should point to real routes for growth. It should name strengths, not just gaps. It should capture where the learner wants to go.
The sixth outcome is the one our tools miss most: areas of continued support. That means splitting two different things. Some items are instructional targets, skills to teach. Others are environmental supports, things you keep in place for life.
A support is not a failure to teach. It is often the faster route to the life the person wants. So ask where accommodations live in your process. In genArete they get their own section of the report overview. Two facts belong with each one. Why does this support exist? Is it working? A support nobody checks is just a habit.
Skill building is not the only option we have when a learner is not engaging in skills that are preventing them from accessing opportunities environments that are important to them.
From the talk — Mark Malady, BCBA
Design Choices Behind the genArete Assessment
Start with cold probes (testing a skill without teaching). The logic says teaching would spoil the data. So the first sessions with a new learner go: do this, you cannot, next item. No feedback. No reinforcement. That is a rough way to meet someone. Teams patch it with a pairing session first, which quietly admits the problem. genArete drops cold probes instead. Careful measurement does the same job. The data show whether a skill sits in acquisition (still being learned) or in mastery.
Entry points are the second choice. Skills sit in clusters. A cluster might hold five items, from a low entry point up to a high one. What you already know about the learner picks where you start. One cluster covers following instructions. The top item is following presumed novel instructions. Below that is following instructions with an instructional history. Below that is acquiring one new instruction-following response with no prior history. That bottom item measures how fast you can teach a single new instruction. We don't want to present the same item to those two learners.
Measurement is matched to the target, not to the tool. Rate, latency, duration, and behavior rating scales all appear across the instrument. Assent (the learner agreeing to take part) is built in through a pre-teach step before each pinpoint (the exact skill being measured). Accommodations are built in too. Learning channels from precision teaching (a rate-based teaching method) set how input and output pair. Timing windows can shrink for learners with low tolerance for instruction, and the measure keeps its power.
The process opens with engage, a 30 to 45 minute semi-structured interview (guided but flexible). It samples every domain. It sets the comparison criteria. It marks which domains are not relevant right now. Results so far: more than 400 learners completed, ages 2 to 68. Four outside organizations have run it, including three adult services. Rehabilitation programs reported a tenfold rise in skills mastered per year. A case comparison and a test-retest study are in peer review.
We don't want to present the same item to those two learners.
From the talk — Mark Malady, BCBA
Milestone Comparison Criteria Instead of Same-Age Norms
Picture a planning meeting for a 30 year old adult. Eight or nine people sit at that table. Run your skill assessment on all of them. Everyone scores 100 percent except the person the meeting is about. That is a ceiling problem in the tool, not a truth about people.
A milestone is a real goal the learner wants. You can build criteria from the people who already do it. Find competent performers who resemble your learner. Benchmark their skills. Set item criteria from the lowest, median, or average performer. Test those criteria on two new performers who share your learners' demographics. Then use it with your learner.
genArete reports on a radial graph (a round spider-web chart) called the skillgram. The outer edge is the terminal target (the highest score possible). A second line marks the milestone boundary. One adult in 24 hour care wanted to stay home alone for 10 to 15 minutes. The team kept adding skills to teach first. The graph compared that adult to people in similar settings who already stay home alone. The adult was already past the line. The question changed from what to teach to what the support team needed to feel comfortable. That adult now stays home for up to five hours.
Other profiles show a skew: strong in some domains, thin in others. The strong domains become the starting point. In one case a family worried about caring for a pet. The learner focused on self-care skills and felt ready. The family focused on other domains and did not. Neither side was wrong, and the graph gave them shared language. You can do a lighter version with tools you already own. If a learner wants ballet class, highlight the VB-MAPP items ballet class needs. Those become the targets. The rest of the boxes can wait.
We think that in allocating your time to one area, you by definition are not allocating your time to something else.
From the talk — Mark Malady, BCBA
Audit Your Assessment Selection Across Your Last 10 Learners
Here is a test you can run this week. Take your last 10 learners. Give each assessment package its own code. One learner got three tools. Another got two. Another got one. Put the codes in a simple pie chart.
Same package for everyone means no real differentiation. You can use several instruments and still treat every learner the same way. That is the trap the multimodal advice is meant to prevent.
Fix it at intake. Use a semi-structured interview to decide which tools to run. Note which ones you can skip because the learner is likely strong there. Note which ones do not fit the learner's current path. Then check whether two assessors would make the same picks from the same intake. That is reliability of selection, and almost nobody measures it. When the instrument leaves materials to your discretion, match them to what the learner actually likes. None of this touches fidelity (running the tool as designed).
If everybody, if you have one color, that should be a red flag.
From the talk — Mark Malady, BCBA
Getting an Emerging Instrument Approved by a Funder
Funders shape this whole picture. They pay for certain tools. Some name the tool you have to run. That is a big part of why legacy instruments keep their market share. It is also why a team that wants to try something newer stalls out.
The short-term move is to run the new tool next to the old one. Take the ask to your reviewer before you need it. Could we run the VB-MAPP plus this instrument, for these learners, in this way? Then be transparent about why. Show them the differentiation chart from your own caseload. Show what you wanted to do and what you felt forced to do. That contrast opens more space than a complaint does. It also helps to know the person on the other end of the phone. Many reviewers do care about this work.
There is a written standard most clinicians never see. Malady checked policies across 15 different insurers. The bar was the same every time: one validity study and one reliability study. If they push back, ask for the policy. Ask for the code that defines evidence-based assessment for that plan. Many payers bury it. They still have to have one.
So if you have one reliability and one validity study, submit that with the first time you submit an instrument.
From the talk — Mark Malady, BCBA
Common questions
Are the VB-MAPP and ABLLS-R evidence based?▾
As of the 2025 review cited in this talk, both sit in the emerging category for reliability and validity. That is not the same as no support, and newer work is testing them on very large datasets. It does mean the tools we use most have thinner validation than some tools we barely use, like the ABLA.
What is the difference between norm-referenced and criterion-referenced assessment?▾
Norm-referenced items come from a comparison group, so the score says how the learner stacks up against same-age peers. Criterion-referenced items come from a logical account of what mastery requires, with a set mastery criterion. Each has limits. Norm-referenced scores ignore individual circumstances, and criterion-referenced tools invite teaching to the test.
How do I individualize assessment without breaking fidelity?▾
Change what you select and how you prepare, not how the items are run. Use a semi-structured intake interview to pick which instruments to run and which to skip. Match any clinician-chosen materials to the learner's known interests. Build milestone overlays that highlight the items tied to the learner's real goal.
What is a cold probe and why would I stop using one?▾
A cold probe tests a skill with no teaching, no feedback, and no reinforcement, so the data stay clean. The cost is that your first interactions with a new learner become a string of failures. Good measurement systems can tell acquisition from mastery without that cost. Teams that drop cold probes also report more use of other standard teaching strategies.
How do I build milestone-based comparison criteria?▾
Start from the goal the learner wants, such as staying home alone or starting kindergarten. Find competent performers who resemble your learner and benchmark their skills. Set item criteria from the lowest, median, or average performer. Then test those criteria on two new performers with shared demographics before you use them.
About the speaker
Mark Malady is a BCBA and an autistic clinician who works on the genArete assessment and runs a skill-based learning center. He has worked in the field since 2007 and has been a BCBA since 2013, across learners from age two to over 70. He names his work on the assessment as a conflict of interest, and he argues that professional tools can cause harm without ongoing ethical scrutiny.
This summary was generated from the recording’s transcript. Quotes are taken word for word from the talk.
What you'll learn
- 1Learning Objectives
- 2By the end of this CEU, participants will be able to:
- 3Describe the current state of empirical support for common skill-based assessments and the disconnect between evidence and field utilization
- 4Identify features of skill-based assessment instruments, such as norm-referenced comparison criteria and mastery-of-all-items graphical displays, where normalization agendas can appear
- 5Explain learner-informed assessment strategies, including semi-structured intake interviews, flexible entry points, and milestone-based comparison criteria, that individualize assessment without compromising instrument fidelity
Concepts in this CEU
Related talks
Talks that cover the same concepts

Generate: Learner Centered Skill Assessments
The Behaviorist Bookclub
genArete: Milestone based comparison criteria in Skill Assessment
The Behaviorist Bookclub

Assent: Don't just say Yes!-
The Behaviorist Bookclub
Delegating with Confidence: Training and Supervising RBTs to Support Assessments
Copper Consulting Group

Getting the Right Answer Isn’t Enough: Establishing Speaker–Listener Repertoires for Complex Learning
The Behaviorist Bookclub
Beyond Compliance: What Assent-Based ABA Looks Like in Real Life
Dani Elle Mentorship and Consulting
More from The Behaviorist Bookclub
Supervising Beyond Technical Skills: Professionalism, Communication, and Clinical Judgment
The Behaviorist Bookclub
Reflective Supervision and Client Outcomes
The Behaviorist Bookclub
When the Foundations Are Right: Unlocking Advanced Verbal Repertoires
The Behaviorist Bookclub
OpenCEU is supported by advertising from third-party sponsors. Sponsors are clearly labeled, have no influence over course content, instructors, or CEU decisions, and never appear on certificates. The BACB does not sponsor, approve, or endorse OpenCEU sponsors or their products.
