genArete: Milestone based comparison criteria in Skill Assessment
Presented by Mark Malady, BCBA
In this second talk of his three-part series on skill-based assessment, Mark Malady, BCBA, examines the comparison criteria embedded in the assessments most commonly used with autistic learners and argues that many still steer clinicians toward normalization even as the field claims to have moved past that agenda. He shows how visual design choices, such as center-to-outer layouts, preset axes, and blank scoring space, evoke predictable clinician behavior like teaching to the test, where assessment scores rise while global measures such as the Vineland show no meaningful change. He frames assessment tools as products of clinical work that acquire evocative and abative functions over clinical decision-making. Malady contrasts the two dominant criterion types, age-based norm-referenced benchmarks and professionally developed criteria, and proposes a third alternative: flexible, person-selected, outcome-based milestones benchmarked from competent performers. Using the genArete Learning System's skill-gram displays, he demonstrates how overlaying a milestone as a second data path redirects clinical attention and changes the meaning of blank space, illustrated by an adult in 24-hour care whose data far exceeded a benchmarked stay-home-alone milestone, reframing a skill-building conversation as a dignity-of-risk conversation. He reviews benchmarked milestones such as entry-level employment, living with a roommate, and kindergarten readiness, then outlines practical steps: identify competent performers, set item criteria, test with novel performers, build cross-learner data flow, and use quality-of-life surveys to repeatedly orient to milestone achievement. The Q&A covers selling the approach to leadership, selecting five focus targets, and reconciling learner and caregiver milestone mismatches.
Register free & watch nowCancel any time. Cert delivered to your inbox the moment you pass the quiz.
Learner ratings
from 39 learners
96% would recommend this CEU to other professionals in their field (24 responses)
“Good content ntl the excessively motonous tone made it super boring.”
— Reagan P.“This presentation made me realize the importance of outcome-based milestones and how they relate to the world of Applied Behavior Analysis. Outcome-based milestones are important since they ensure therapy translates into real-world, meaningful skills rather than just compliance in a controlled setting.”
— Cristina G.
About this CEU
In this second talk of his three-part series on skill-based assessment, Mark Malady, BCBA, examines the comparison criteria embedded in the assessments most commonly used with autistic learners and argues that many still steer clinicians toward normalization even as the field claims to have moved past that agenda. He shows how visual design choices, such as center-to-outer layouts, preset axes, and blank scoring space, evoke predictable clinician behavior like teaching to the test, where assessment scores rise while global measures such as the Vineland show no meaningful change. He frames assessment tools as products of clinical work that acquire evocative and abative functions over clinical decision-making. Malady contrasts the two dominant criterion types, age-based norm-referenced benchmarks and professionally developed criteria, and proposes a third alternative: flexible, person-selected, outcome-based milestones benchmarked from competent performers. Using the genArete Learning System's skill-gram displays, he demonstrates how overlaying a milestone as a second data path redirects clinical attention and changes the meaning of blank space, illustrated by an adult in 24-hour care whose data far exceeded a benchmarked stay-home-alone milestone, reframing a skill-building conversation as a dignity-of-risk conversation. He reviews benchmarked milestones such as entry-level employment, living with a roommate, and kindergarten readiness, then outlines practical steps: identify competent performers, set item criteria, test with novel performers, build cross-learner data flow, and use quality-of-life surveys to repeatedly orient to milestone achievement. The Q&A covers selling the approach to leadership, selecting five focus targets, and reconciling learner and caregiver milestone mismatches.
From the talk
What was covered
Mark Malady, BCBA on why skill assessments still nudge toward normalization, and how a milestone overlay changes what you teach.
- Notice where your assessment display pulls your eye before you write goals. The page is already making part of that choice for you.
- If assessment scores climb while broad life measures stay flat, suspect teaching to the test.
- You can add a milestone as a second data path on top of the assessment you already own.
- Benchmark a milestone by scoring people who already do the thing. Then set item criteria from their lowest, median, or average score.
- Run quality of life check-ins every three to six months. Ask what the learner actually got to do.
- When a learner and a caregiver want different milestones, overlay both and teach the skills that serve both.
Why Skill Assessments Still Steer Toward Normalization
Behavior analysts have spent years answering hard questions from the autistic community. Much of the answer sounds the same. The field says it has moved past normalization (making learners look non-disabled). Mark Malady is not sure the tools agree. He argues that many common assessments still point us that way. This is not a claim about anyone's intent. It is a claim about what the instrument makes easy to do.
He grounds the argument in basic science. Our subject matter is one person in one context. That sits inside the operant (a behavior shaped by its results). It is also why our research base was built on single case designs (studies of one learner at a time). Malady says that history sits oddly next to tools built on group norms.
He points to one paper he likes. It is called Community views of neurodiversity models of disability and autism intervention, from Patrick Dwyer and colleagues. It came out online in 2024 and is often cited as 2025. His read is that people on both sides of the social and medical model divide agree on one thing. Individualized support that meets the learner's context is good.
Then he turns the science on us. Our tools are products of clinical work. They can pick up evocative and abated functions (cues that trigger or reduce behavior). A score sheet can prompt one recommendation and quiet another. That is how a tool can walk a careful clinician toward indistinguishability (looking no different from peers).
And clinicians are fundamentally behaving organisms, which means we are underneath all of the same rules that we learn about our science.
From the talk — Mark Malady, BCBA
How Blank Space on a Score Sheet Picks Your Next Goal
Malady put a common assessment display on screen. It is a set of rings. Skill categories run from the center outward. A few boxes are filled. Most are blank. Then he asked a simple question. Where does your eye go?
Most clinicians land on the empty categories in the bottom ring. Three of them were completely blank. That looks like the best return on effort. His point is that the layout made that call, not the data. The page told us where to start.
He tied this to a pattern many of us have watched. A new clinician runs an assessment. They list every missing skill. They teach that list, sometimes using the assessment items themselves. Six months later the scores are higher. Then someone asks how life is going. Broad measures like the Vineland (a wide daily living skills scale) show no real change. Support network members report nothing new.
Why does this happen to people trained in generalization (using a skill in new situations)? Malady thinks the answer is almost too simple. A center to outer layout makes a line. We are verbal organisms with long histories of finishing lines. Preset axes make it worse. When you cannot change the categories, your next move gets very predictable.
Filling the space becomes the goal.
From the talk — Mark Malady, BCBA
Age Norms, Professional Criteria, and a Third Option
Malady named the two criterion types that dominate our assessments. The first is norm referenced (scored against a group average). You learn how a whole group performs on an item. Then you place your learner against that group. In our field the group is almost always an age. Everyone who is three. Everyone who is five.
Norm referencing is not a bad idea by itself, and he said so plainly. Outside our field the group can be second graders, or people working a certain job. The trouble is which norm we picked. We benchmarked against time on earth. We did not benchmark against anything the learner wanted.
The second type is criterion based (scored against a fixed standard). Each target carries a performance criterion. Meet it and the skill counts as mastered. That gets closer to useful. The gap shows up one level higher. Who were the people used to set that criterion? What did their performance let them actually do?
So he offers a third option to sit beside the other two. Use flexible, person-selected, outcome-based milestones. Benchmark them from people who already do the thing. Any instrument can carry several of them. You pick the one that matches the person in front of you.
Right now, most of them are extremely vague or ambiguous on where the criterion came from or what its purpose was.
From the talk — Mark Malady, BCBA
Overlaying a Milestone as a Second Data Path
The display he uses is a radar chart. His team calls it a skill-gram (a spoke chart of skill areas). Each spoke is a domain. Each domain converts to one shared scale. That shared scale is what makes the next step work. You can lay a second line over the learner's data. That second line is the milestone.
He walked through one case from his own clinical work. An adult living in 24-hour care wanted to stay home alone. State law allowed it. The service model never made it possible. The goal had been on the plan for years. Everyone said it was a fine goal. Something always came up when it got close.
So the team benchmarked it. They asked what it really takes to stay home alone without risk. Then they drew that milestone on the chart. The learner's data ran well past it. The skills were already there. The opportunity was not.
That changes what the meeting is about. It stops being a teaching plan. It becomes a question of dignity at risk (the right to take a chance). Malady also points at the gap itself. The distance between the two lines carries information. It supports relative decisions instead of a ranked deficit list.
So it changed it from a skill-building conversation to a dignity at risk conversation.
From the talk — Mark Malady, BCBA
Same Data, Two Milestones, Two Different Programs
In the standard display, the outer edge is the milestone. Filling the page becomes the job. A second data path breaks that hold. The milestone now sits wherever the benchmark puts it. Blank space past that point stops driving the plan.
He showed one learner's data twice. The first overlay was an entry level employment milestone. Attention pulled away from the outer edge right away. Committed action skills (skills for following through on goals) sat above what the job asked for. Complex verbal skills (advanced language and communication skills) looked strong. Movement was a little short. Pivotal skills (skills that unlock broader learning) and self care were close. Community contact was far away.
Then he swapped the milestone and kept the data. The second one was about using values and goals to drive change. That is a skill set, not a mood. Movement mattered less here. Complex verbal mattered more.
He made the reason concrete. Someone in a first job may not know how to move up. They may not know what fair or unfair treatment looks like. They may not know how to work with a supervisor. All of that is complex verbal work. Same learner, same scores, different plan.
None of us can do everything.
From the talk — Mark Malady, BCBA
The Milestone Library He Has Benchmarked So Far
Malady listed what his team has benchmarked and what is still open. Staying home alone is done. So is entry level employment, which he described as jobs near minimum wage. Entry to trade school and a first successful year is done. Kindergarten and fifth grade are done. Nothing above fifth grade is benchmarked yet.
Several others are finished on the learning side. Emergent communicator covers early sharing of wants and needs. Directed learner with one on one support covers learning from a paraprofessional or RBT (Registered Behavior Technician). Peer group without supervision covers hanging out with same age peers and enjoying it. Others cover tolerating another person's decision and talking a problem through to resolution. More cover team sports and solo sports like gymnastics or martial arts.
Some are still in progress. Living with a roommate is close but not finished. Living with a significant other is not built yet, and he wants it. Dating apps is one he is eager to benchmark. Support networks often react with fear instead of a plan. He also described a career rung climber milestone for moving past a first job.
The point behind the list is a challenge to how we use the word independence. He keeps asking what the independence is for. Living alone. Choosing a job. Staying safe with other people. Building a real relationship.
It's independence for what?
From the talk — Mark Malady, BCBA
How to Benchmark a Milestone With the Tool You Already Use
The method is short to write and hard to do. Find competent performers, meaning people who already do the thing well. Score their skills on your instrument. Set the criterion for each item from that group. You can use the lowest score, the median, or the average. Malady leaves that call to you.
He adds a bonus step that is worth the trouble. Test the draft milestone on two new competent performers who share demographics with your learners. If it holds, teach your learner to that criterion and see what happens in the real setting.
He showed an example from a colleague at UNR. They built a ballet class milestone using the VB-MAPP (a common early skills assessment). The learner's data was one line and the milestone was a second line. The three boxes on the bottom that would normally grab a clinician's eye turned out to be irrelevant. Those skills were not what the ballet class asked for.
Malady also warns that milestones move. A milestone that matters this year may not matter in three years. That is not a flaw in the method. It matches how people actually live, and it means your criteria stay in motion too.
So we can overlay a milestone on top of any assessment.
From the talk — Mark Malady, BCBA
Data Flow, Quality of Life, and the Two Hard Conversations
Benchmarks come from your own caseload, so your data has to move. Most clinics tie one learner's data together well, especially under insurance. What usually does not exist is flow across learners, and across clinicians in the same organization. Malady suggests adding one filter: what new thing happened in this person's life this month. A birthday party. A first sleepover. Then connect that answer back to performance and assessment data. He noted this is easy in a spreadsheet and messier in a centralized system.
Asked how to sell this to leadership, he pushed back on the framing of extra work. Start with what already exists. Case notes usually record the big life events families report. The real question is how you extract them. He would first look at whether the team already reviews mastery across learners, since that infrastructure is the entry point.
For quality of life, he uses a well-being scale plus an intake and outtake survey. Both ask the same thing in two directions. Tell us something the support person did that mattered. Tell us something the child got to experience. He recommends collecting this every three to six months.
Two practical questions closed the talk. When a learner has needs everywhere, his team filters forty to sixty pinpoints (single countable target skills) down to about fifteen. The family and the learner pick from there. When the learner and the caregiver want different milestones, they run a motivational interview (a guided conversation about goals) with both. Then they overlay both milestones. A teen who wants the skate park and a parent who wants church share more skills than either expects. Naming the overlap is what starts the compromise.
So we really take it from like a choice or a menu perspective.
From the talk — Mark Malady, BCBA
Common questions
What is a milestone-based comparison criterion in skill assessment?▾
It is a benchmark built from people who already do the thing your learner wants to do. You score those competent performers on your assessment. Then set an item criterion from their lowest, median, or average performance. That benchmark becomes a second line on the learner's chart. The comparison is to a real outcome, not to an age.
How is a milestone different from an age-based norm?▾
An age norm compares a learner to everyone who has been alive the same number of years. A milestone compares them to people who already do a specific thing. That could be holding an entry level job or staying home alone safely. Malady argues the age norm quietly sets normalization as the goal. The milestone sets the learner's own stated outcome as the goal.
Do I need a special assessment tool to use milestones?▾
No. Malady says a milestone can be overlaid on any assessment, including tools you already own. What you need is access to enough competent performers to score. You also need a way to put both data sets on one shared scale. His own example used a ballet class milestone built on the VB-MAPP.
What does teaching to the test look like in skill assessment?▾
A clinician runs the assessment, lists every unscored item, and teaches that list directly. Scores rise on the next administration. Broad measures like the Vineland show no change, and families report nothing new in daily life. Malady argues the assessment layout invites this, since blank space reads like a to-do list.
What if the learner and the caregiver want different milestones?▾
Malady's team interviews both parties separately, then overlays both milestones on the same skill data. Shared skill gaps show up visually. Those overlapping skills become the compromise, and the family decides how to sequence the rest. Naming the mismatch out loud is often what unlocks movement.
About the speaker
Mark Malady, BCBA has 18 years of clinical experience running skill-based assessments across the lifespan. He has worked with learners holding a wide range of diagnoses and support needs. He is an autistic adult. He discloses a level one PDA profile (a strong drive to avoid everyday demands) and a twice-exceptional school identification (gifted and having a learning difference). He still works clinically in a short-term, low-dose service model. He states his conflict of interest openly. He created the skill-based assessment discussed in this talk, and the product is owned by his employer.
This summary was generated from the recording’s transcript. Quotes are taken word for word from the talk.
What you'll learn
- 1Learning Objectives
- 2Participants will be able to label some common comparison criteria used by primary skill-based assessments within ABA based services for autistic learners.
- 3Participants will be able to identify the benefits of milestone-based comparison criteria and the risk of using the criteria.
- 4Participants will be able to describe the relationship between indistinguishability and norm-referenced criterion and strategies that can be employed to decrease this risk.
Concepts in this CEU
Related talks
Talks that cover the same concepts
genArete: Learner-Centered Skill Assessment
The Behaviorist Bookclub

Generate: Learner Centered Skill Assessments
The Behaviorist Bookclub
Delegating with Confidence: Training and Supervising RBTs to Support Assessments
Copper Consulting Group

Getting the Right Answer Isn’t Enough: Establishing Speaker–Listener Repertoires for Complex Learning
The Behaviorist Bookclub
Child Development for BCBAs- Age 9-11
The Behaviorist Bookclub
Child Development Deep Dive: Middle Childhood (6-8 year olds)
The Behaviorist Bookclub
More from The Behaviorist Bookclub
Supervising Beyond Technical Skills: Professionalism, Communication, and Clinical Judgment
The Behaviorist Bookclub
Reflective Supervision and Client Outcomes
The Behaviorist Bookclub
When the Foundations Are Right: Unlocking Advanced Verbal Repertoires
The Behaviorist Bookclub
OpenCEU is supported by advertising from third-party sponsors. Sponsors are clearly labeled, have no influence over course content, instructors, or CEU decisions, and never appear on certificates. The BACB does not sponsor, approve, or endorse OpenCEU sponsors or their products.
