REAL School BudapestWellbeing

How we built it

Why a competency-based approach

The thinking behind how learning is structured and assessed in this programme.

Written primarily for teachers and curriculum colleagues; suitable for accreditors and peer schools. This framework was developed at REAL School Budapest — a small independent school operating with substantial teacher autonomy — in sustained collaboration with AI as a co-intelligence. The agentic method page describes how that collaboration was structured and what implementation context the framework assumes.

What’s wrong with grades

Imagine your child comes home with a B in Spanish. What does that tell you? Can they introduce themselves? Order food? Ask a stranger for directions? Hold a conversation about something they care about? A single letter can’t answer any of these questions. It collapses a developing person into a number compared against an average — and in doing so, it hides exactly the information a learner needs to grow.

Aggregated grades have always been a poor signal of what someone can actually do. Employers have known this for decades; assessment systems are slowly catching up.1 In the meantime, schools that want to take learning seriously have to build something better.

This is true for every subject, but it’s especially true for wellbeing. A grade for “self-awareness” or “resilience” wouldn’t just be uninformative — it would be a category error. We don’t grade who children are. We help them see themselves more clearly, and we describe what we observe in language a young person can actually use. That requires a different kind of structure.

The shape of the framework

The programme organises wellbeing learning across eight competencies: emotional and social self; attention and reflective practices; physical wellbeing and self-care; consent, safety and healthy relationships; community, purpose and belonging; wellbeing science; metacognition and self-regulated learning; and critical digital literacy.

Each competency contains two to five learning targets — specific things a student is learning to do. Each learning target carries a set of criteria describing what quality of performance looks like at five levels: No Evidence, Emerging, Developing, Competent, and Extending. You can browse all of this in the curriculum explorer.

That structure — competency, learning target, criterion, level — is the skeleton. The rest of this page explains the thinking behind it. What does a competency actually contain? How does competence develop? What’s the evidence for this approach, and where does it run out? Getting the structure right depends on getting the conceptual grounding right first.

What does a wellbeing competency consist of?

Before getting into the structure, it’s worth being precise about what a competency actually is — because the term gets used loosely, and the looseness causes problems downstream.

A competency, in the sense used by most contemporary frameworks (the EU’s GreenComp, LifeComp, EntreComp; the OECD Learning Compass 2030), is the integrated capacity to mobilise knowledge, skills, and dispositions to act effectively in a real-world situation. The three components do different work:

Knowledge is propositional content — claims a learner can articulate. The structure of the nervous system. The distinction between guilt and shame. The mechanics of consent. This is what philosophers call knowing-that.

Skills are procedural — the ability to do something with reliability and quality. Performing a breathing technique. Naming an emotion accurately. Holding a difficult conversation. This is knowing-how: knowledge in performable form.

Skills and knowledge aren’t separate categories. Most cognitive scientists treat skills as procedural knowledge — knowledge that has been compiled and made performable. The two are different ways the same underlying capacity is held and expressed.

Dispositions are stable orientations toward the world. Curiosity. Honesty. Self-regulation. Empathy. They include attention, valuing, judgement, and inclination, not just behaviour. A curious person isn’t just someone who asks questions — they’re someone who notices the world differently, values inquiry, and is willing to sit with not-knowing.

Most teachers will recognise the lived shape of dispositional development from their own classrooms — the slow shift in how a young person engages, the moments when something integrates that hadn’t before, the way you notice a student is somehow different in a way that’s hard to name precisely. Dispositions don’t develop on a clean timeline. They emerge across years, through hundreds of small moments, and they consolidate when the learner is in a relational environment that supports them.

A real competency integrates all three. Take self-regulation under stress. The competent learner has knowledge (they understand what nervous-system states are, why regulation matters), skills (they can perform a regulation strategy), and dispositions (across many situations, they tend to notice when they need to regulate, value doing it, and act on it). These aren’t three separable things the learner experiences as separate. They develop together through repeated lived experience in real situations.

Yoga is a useful illustration. Performing an asana is a skill — procedural knowledge expressed in the body. Understanding why a particular pose works the body in a certain way is propositional knowledge. The orientation toward attention, presence, and embodied awareness that develops through the practice is dispositional. A practitioner doesn’t first learn the asanas, then add attention. The attention is part of how the asana is learned. Wellbeing competencies have the same structure.

This matters for how we design and assess. The skill and knowledge components can be assessed through performance against criteria. The dispositional component cannot — not reliably, anyway, because the orientation lives in the relationship between the learner and the situation, not purely in the learner. A student who can self-regulate in a calm classroom may not be able to in conflict with a sibling. A student who is honest in low-stakes settings may struggle when honesty has costs. Trying to grade dispositions on a rubric produces unreliable signals at best, and damages the relational ground the programme depends on at worst.

So the framework’s design choice is to treat the three components differently in assessment, while recognising that the learner experiences them as one.

One further distinction matters here, because the term gets used loosely in wellbeing contexts. A behavioural habit is an automatic action triggered by a context, developed through repetition — brushing teeth before bed, doing breathing practice when the timer goes off. Habit formation is one of the most evidence-supported ways to change behaviour over time, and behavioural-habit work is part of the wellbeing programme.

A behavioural habit is not the same as a disposition. A student can mechanically perform three breaths when frustrated without actually orienting toward regulated response. The habit alone doesn’t constitute the disposition; the disposition is the orientation that the habit can support. Habits are tools for dispositional development — they create the lived practice within which dispositions consolidate. But the habit alone is not the goal.

How competence develops

Think of something you’re good at — driving, cooking a familiar dish, reading a room. There was a time you couldn’t do it at all, and didn’t yet know what doing it would involve. Then a time you could just about do it, with full concentration. Now you mostly do it without thinking. That felt experience of getting good at something — from invisible, to effortful, to automatic — is the most intuitive way into how competence develops, and it was described five decades ago as the four stages of competence (Broadwell, 1969):

  1. Unconscious incompetence — the learner doesn’t know what they don’t know.
  2. Conscious incompetence — the learner is aware of what they need to learn and starting to see how.
  3. Conscious competence — the learner can perform the skill deliberately, with attention.
  4. Unconscious competence — the learner performs the skill fluently, without conscious effort.

The shift from unconscious to conscious incompetence is the development of awareness — and awareness is the prerequisite for everything that follows. A learner who doesn’t know they don’t know cannot work toward knowing; the territory itself is invisible to them. A clearly named target — “recognising and naming my emotions” — makes that territory visible. Knowing what we don’t yet know is what allows us to grow: to become more skilled, more capable, more able to act on the world. This is one of the underrated benefits of a competency-based approach. It makes the developmental destination explicit, for teachers and students both, so that attention can land in the right place.

Broadwell's four stages of awarenessA single left-to-right progression of four boxes: Unconscious incompetence, Conscious incompetence, Conscious competence, Unconscious competence. An arrow beneath runs from “Doesn't know the skill exists” on the left to “Does it without thinking” on the right. The figure shows one skill in one context becoming automatic over time.Broadwell — four stages of awarenessUnconsciousincompetenceConsciousincompetenceConsciouscompetenceUnconsciouscompetenceDoesn't know the skill existsDoes it without thinking
Broadwell’s four stages of awareness — a genuine sequence, for one skill in one context.

Broadwell’s arc really is a sequence: for one skill, in one context, awareness comes first and automaticity comes last. That makes it the right way into this territory. It also sets up the one distinction this page most needs to keep sharp. Broadwell’s “competence” describes the learner’s automaticity — how much conscious effort the skill still costs. The level this framework calls Competent describes something else entirely: the quality of a piece of work, judged against criteria. The same word, doing two different jobs, in two different lenses. The rest of this page keeps them apart.

Three lenses on developing competence

“Developing competence” sounds like one thing. It isn’t. There are at least three questions inside it, and the frameworks educators reach for answer one question each. This framework treats them as three lenses. Each sees something the other two are blind to — and none of them maps onto another.

The first lens is the kind of thing being learned. Multiplication facts, historical interpretation, and being a good friend do not develop the same way, and a framework that treats them the same will fail at least one of them. This framework distinguishes three patterns — hierarchical content, horizontal content, and dispositional content — taken up in detail later on this page.

The second lens is the quality of a specific performance, right now. This is what the five levels are for. A level is a judgment about one piece of work at one moment — this piece of writing, this conversation, this demonstration — held against criteria. Emerging and Developing describe work that partly meets the criteria. Competent describes work that meets them. The criterion lives at Competent: it is the target itself, not a stage on the way to somewhere else. Extending describes depth visible within the same performance — work that goes beyond the criteria in sophistication: justifying choices, anticipating counterarguments, handling edge cases, connecting the work to adjacent ideas. Extending is not a claim about what the student can do elsewhere. It is more depth in this piece of work, in this context. (No Evidence sits below all of these. It indicates a student hasn’t yet had the opportunity to demonstrate the skill, not that they can’t.)

The five levels as a quality scaleA vertical scale of five boxes, from No evidence at the bottom up through Emerging, Developing, Competent, and Extending at the top. An upward arrow on the left is labelled “Quality of this performance”. Annotations on the right mark Competent as the criterion — work that meets it; Extending as optional depth beyond the criteria within the same piece of work; and No evidence as no opportunity yet, not inability.The five levels — a quality scale for one performanceQuality of this performanceNo evidenceEmergingDevelopingCompetentExtendingThe criterion —work that meets itOptional depth beyond the criteria,visible in the same piece of work —not a claim about new contextsNo opportunity yet —not inability
The five levels — a quality scale, applied to one performance at one moment.

This is why criteria at different levels must describe genuinely different qualities of work, not the same skill in increasingly enthusiastic language. Contemporary research keeps reinforcing the principle underneath — learners get better at the cognitive activity they actually rehearse. A pair of recent studies, synthesised by Hendrick (2026), found that students who practised remembering rules got better at remembering rules but no better at applying them.2 When we write a level descriptor, it specifies the kind of quality visible in the work at that level — and quality is all it specifies. The conditions the work was produced under belong to a different lens.

The third lens asks a question no single performance can answer: under what conditions can the learner do this reliably — over time, and across situations? This is the question the instructional hierarchy (Haring, Lovitt, Eaton & Hansen, 1978) was built to answer. It names four stages:

  1. Acquisition — building the skill: starting to perform it but not yet with accuracy or consistency.
  2. Fluency — accurate and faster, with reduced cognitive effort.
  3. Generalization — accurate and fluent across different contexts and situations.
  4. Adaptation — adapting the skill to genuinely new situations.

Notice what kind of claims these are. None of them is a judgment about one piece of work. They are patterns across many occasions: how easily the skill comes, whether it survives a change of setting, whether it can be bent to fit a situation nobody prepared the learner for. That is also why the hierarchy’s proper home is instruction rather than scoring — it tells a teacher what kind of practice a student needs next: more accuracy-building, more repetition until the skill comes easily, more varied contexts, more open problems.

The instructional hierarchy — conditions over timeA single left-to-right progression of four boxes: Acquisition, Fluency, Generalization, Adaptation. An arrow beneath runs from “First attempts” on the left to “Genuinely new situations” on the right. The figure describes the conditions under which a learner can perform reliably, over many occasions and across situations — not the quality of any single performance.The instructional hierarchy — conditions over timeAcquisitionFluencyGeneralizationAdaptationFirst attemptsGenuinely new situations
The instructional hierarchy — a lens on conditions over time. Its own progression, not a set of rubric zones.

Where the lenses meet — and where they come apart

The rubric level and the hierarchy stage answer different questions. “How good is this work?” is not the same question as “under what conditions can they do it?” — and answering one carefully will never answer the other.

Within a single, familiar context, the two rise together. A student acquiring a skill in one classroom, on one kind of task, will typically produce Emerging work, then Developing work, then work that meets the criteria. Quality climbs as acquisition proceeds. This is why the aligned picture feels true: watch one student in one context, and the rubric looks like a staircase up the hierarchy.

Across contexts, they come apart. A student who has genuinely generalised a skill — who can use it in several familiar settings — can produce lower-quality work the first time they face a genuinely new context, and then climb again as they adapt. That dip is not a regression to acquisition. It is information about generalization: it shows exactly where the skill’s reach currently ends.

So the rule this framework runs on: the two lenses coincide within a context and diverge across contexts. That is why the five levels cannot be laid over the hierarchy as corresponding rows. Any picture that aligns them takes the within-context case and silently claims it holds everywhere — and across contexts, exactly where assessment matters most, it doesn’t hold.

One further piece of evidence that these are different instruments. Fluency — producing the same quality of work with ease, quickly, at lower cognitive cost — is essentially invisible to a quality rubric. Same quality, less effort: the rubric records nothing. A whole stage of the hierarchy passes without leaving a trace on the quality scale. That is not a flaw in the rubric. It is proof that the two measure different things.

Two traps this framework refuses

Fusing the lenses produces two specific mistakes. Both are common enough in competency frameworks to be worth naming — and an earlier version of this page committed both, so we name them with some humility.

Trap 1: mapping rubric levels onto hierarchy stages. The tempting move is to line the frameworks up — Emerging means acquisition, Developing means late acquisition into fluency, Competent means fluent performance, Extending means generalization and adaptation. An earlier version of this page made exactly this claim, and drew the three frameworks as aligned rows on a single axis. The coincide-and-diverge rule is what unpicks it: the alignment holds only while you watch one context. A rubric level is a reading at one moment; a hierarchy stage is a pattern across many moments and many contexts. A ladder cannot represent a matrix — and drawing them as aligned rows asserts that it can.

Trap 2: writing condition-claims into the top level. The earlier version of this page defined Extending as “applying the skill in novel situations, teaching others, integrating it with other skills.” Every clause of that is a claim about conditions — new situations, other people, other skills — smuggled into what is supposed to be a quality judgment of a single performance. The corrected definition is the one this page now uses: Extending is depth visible within a single performance in the same context — beyond-criterion sophistication in one piece of work. Applying a skill in new contexts and teaching it to others are real, and they matter — but they are signs of generalization and adaptation, which belong to the instructional hierarchy, not to any rubric level. Teaching others, in particular, is an instructional move: one of the best ways a teacher can generate strong evidence of generalization. It is not a level a student reaches.

One skill, three contexts

Here is the whole architecture at work on a single skill: constructing a persuasive argument.

Through the first lens, this is horizontal content — one of the three patterns described later on this page. There is no single correct answer a more advanced student reaches; there are better and worse arguments, more and less careful handling of evidence, more and less awareness of the other side. The criteria describe quality of reasoning.

Context A — a class debate on a familiar topic. The first attempts are Emerging: a claim, some conviction, no evidence connecting the two. With teaching and practice the work climbs — Developing, then Competent: a clear claim, supported by relevant evidence, addressing at least one opposing view. With more practice comes Extending, still within the debate: the student anticipates the strongest counterargument before anyone raises it, concedes a weak point strategically, connects the argument to an idea from another subject. Each of those is a rubric judgment of one performance. The climb across them is what acquisition looks like from the outside.

Context B — a persuasive op-ed on an unfamiliar topic. The first attempt drops to Developing. The student can argue — that hasn’t vanished — but the new form and the new content throw them: paragraphs instead of rebuttals, a reader instead of an audience, evidence they had to find rather than evidence they already knew. Over the following attempts they climb back. The dip was not a return to acquisition. It was the exact measure of how far the skill had generalised — the single most informative data point in this whole story.

Context C — a live pitch. The first attempt dips only slightly, to Competent, and recovers fast. The skill is becoming portable: each new context costs less than the last.

Now the failure this example exists to expose. Assess once, at the peak of Context A, and you record Extending and tick the skill off. On paper, the student has mastered persuasion. What that single assessment cannot see is that they could not yet do it in a new context. Real, transferable competence shows up as the rising floor: the quality of the first attempt in each new context climbs — Emerging in the debate, Developing in the op-ed, Competent in the pitch. That rising floor is generalization, and it is visible only by assessing the same competency repeatedly, across contexts, over time.

One skill, three contexts: quality over occasions of assessmentA line chart. The horizontal axis shows occasions of assessment over time, grouped into three labelled context bands: a class debate, a written op-ed, and a live pitch. The vertical axis shows the five levels, from No evidence at the bottom to Extending at the top. A solid performance line climbs from Emerging to Extending during the class debate, drops to Developing when the written op-ed begins, climbs back to Extending, dips only to Competent at the start of the live pitch, and recovers. A dashed floor line connects the first attempt in each context — Emerging, then Developing, then Competent. The floor rises across the three contexts: that rising floor is generalization, which no single assessment can see.One skill, assessed over time, across three contextsPerformance — the quality of each occasion's workThe floor — first attempt in each new context; its rise is generalizationExtendingCompetentDevelopingEmergingNo evidenceFirst attempts rise: Emerging → Developing → CompetentContext AClass debateContext BWritten op-edContext CLive pitch
Within each context, quality climbs; at each new context it drops and recovers. The dashed floor — the first attempt in each new context — rises across contexts. That rising floor is generalization, and no single assessment can see it.

This is the problem of transfer — the thing education most wants and most rarely measures, precisely because assessment tends to happen once, at the peak of a familiar context, before everyone moves on. A framework built to see the rising floor has to be built to look more than once. That is what the criteria, the levels, and repeated observation across contexts are for.

The framework in detail: criteria and levels

The structure was introduced earlier — eight competencies, twenty-one learning targets, criteria at five levels — and the sections above put the levels in their place: a quality scale, one of three lenses. Now to make it concrete: what does a criterion at each level actually look like, for a specific learning target?

The five levels — No Evidence, Emerging, Developing, Competent, Extending — are quality judgments of a specific performance against a specific learning target. They aren’t grades, and they don’t describe the student in general terms. A level belongs to a piece of work, not to a child.

Here’s what this looks like in practice:

The two descriptors are about the same skill — a student naming an emotion they’re feeling. What changes between them is what the student can do independently, what level of distinction they can make, and what quality of self-observation is involved.

This matters for a reason that’s easy to miss. Vague learning targets — the kind that show up in many curriculum documents as one-line outcomes — function in practice to lower expectations rather than raise them. An outcome like “students will understand emotional regulation” is met when anyone says anything about emotions. Without specific criteria for what good looks like at each developmental level, teachers and students drift toward the easiest reading of the target. A clear, criterion-referenced framework is more demanding than a vague one — and also more humane, because students can see exactly what they’re trying to do and where they are.3

One useful tool for putting this kind of self-assessment into practice is the single-point rubric. It shifts more of the responsibility for noticing onto the student — which is part of the developmental work this framework is trying to support — while remaining usable by teachers when they want to set a target and invite a student to reflect against it.

A single-point rubric for self-assessmentA three-column layout showing how a single-point rubric works. The centre column describes the criteria for success, set by the teacher in advance. The left column has space for the student to note areas for growth — criteria they haven't yet met. The right column has space for the student to note areas of competence — criteria they have met, including any they've extended.A single-point rubric — for self-assessment, not gradingAreas for growthCriteria I haven'tyet met(student writes)TargetThe criteria for success(set in advance)Areas of competenceCriteria I have met,including any I'veextended(student writes)
A single-point rubric: the teacher sets the target, the student does the noticing.

In the centre, the criteria for success — set in advance by the teacher, or co-set with the student. On the left, the criteria they haven’t yet met — areas for growth. On the right, the criteria they have met, including any work that extends beyond the target — areas of competence. The teacher can use the same rubric, but its real power is that it asks the student to do the noticing themselves.

Three patterns of how content develops

Different kinds of learning content develop differently — and a framework that treats them all the same way will fail at least one. This section names three patterns the framework distinguishes, and explains how each one shapes how criteria are written and what kind of evidence counts.

Two of the patterns are about how content itself is structured. The third is about how stable orientations of the learner emerge across contexts. The most important practical consequence is in assessment: the first two work with five-level descriptors of performance; the third doesn’t, and needs a different shape entirely.

T1: Hierarchical content. Some content develops in a clear sequence. Multiplication facts. The parts of a cell. The rules of grammar. Spelling conventions. There’s an order, and at each stage it’s possible to specify what mastery looks like. A student who can multiply two-digit numbers can also multiply one-digit numbers, and you can write criteria that capture each step on the path. Five-level descriptors work well here. The sequence belongs to the content — one-digit multiplication before two-digit — while each level still judges the quality of this performance, at whichever step the student is working on.

T2: Horizontal content. Some content develops not by mastering a sequence but by reasoning across multiple legitimate perspectives. Historical interpretation. Ethical analysis. Evaluating competing claims about wellbeing. There isn’t a single correct answer that more advanced students reach — there are better and worse arguments, more and less careful reasoning, more and less awareness of the perspectives that shape a question. Five-level descriptors still work for this kind of content, but they capture quality of reasoning rather than correctness of recall.

T3: Dispositional content. Some of what we want for our students isn’t a skill they perform on demand — it’s a pattern in how they approach the world. Being a good friend. Self-regulating under stress. Showing up consistently for their community. Being honest with themselves. These develop differently from skills, and they need to be assessed differently.

T3 is not a third type of knowledge parallel to T1 and T2. T1 and T2 are about how content is structured. Dispositions are about the learner’s stable orientation, and they cut across content domains rather than sitting parallel to them. A student can be curious about hierarchical content (maths) or curious about horizontal content (history). The disposition is the same; the content varies.

This is why dispositions need a different assessment approach. A student who can self-regulate at school may not be able to at home. A student who is a good friend in a calm classroom may struggle in a high-stakes social situation. A student who shows perseverance with maths may give up immediately on writing.

Trying to grade dispositions on a five-level scale would produce unreliable signals at best — the student’s performance is more about the situation than about a stable level of ability. At worst it would damage the trust the programme depends on.

So for dispositions, the framework uses a different shape entirely. Instead of competency-level descriptors, it uses observation indicators: things teachers and students might notice that suggest a disposition is developing. Each indicator comes paired with confusable behaviours — things that look similar at the surface but mean something different — and conversation prompts that turn observations into developmental dialogue.

This isn’t a grade. It’s an observation that becomes the start of a conversation. Over time, across many situations, with multiple adults noticing, a richer picture emerges of what this student is genuinely working with.

The work of supporting dispositional development is fundamentally relational. Skills you can teach in lessons. Dispositions you mostly model, observe, and discuss. The teacher who wants to support a student’s self-regulation has to be willing to talk about her own. The classroom where a student can practise being a good friend has to be a classroom where being a good friend is what the adults do.

This is why the framework is part of a larger programme. The competency framework is what we measure the educational work against. The relational programme — Daily Community Circle, Reflection 360, Sit Spots, Solo-time, the practices documented on the How it works page — is the substrate within which dispositional development happens. The framework alone isn’t enough. The framework with the relational ground is what makes wellbeing actually grow.

What the evidence says about developing dispositions

The literature on school-based dispositional and social-emotional programmes has matured in ways that should temper expectations. The most rigorous current synthesis — Cipriano et al. (2023), 424 studies, 575,361 students, pre-registered and bias-corrected — finds effects on social-emotional skills around g ≈ 0.22 and on academic achievement around 0.11, roughly half what the foundational Durlak et al. (2011) meta-analysis reported. Several flagship findings have failed or substantially deflated under replication: the long-term predictive power of the marshmallow test (Watts, Duncan & Quan 2018; Sperber et al. 2024); average mindset effects (Macnamara & Burgoyne 2023); grit (Credé, Tynan & Harms 2017); universal adolescent mindfulness (Kuyken et al.’s MYRIAD trial 2022, with adverse effects in some at-risk subgroups). The defensible expectation for a well-implemented programme is small but real and cumulating effects, not transformation. We name this directly rather than overclaiming.

What the evidence does point to with consistent positive signal is a specific set of mechanisms. Direct strategy instruction in metacognition and self-regulated learning is the strongest — Dignath and Büttner (2008) at d ≈ 0.69, Donker et al. (2014) replicated at d ≈ 0.66, and the EEF’s 2025 toolkit rating metacognition and self-regulation at +8 months progress with strong evidence. The active ingredient is combined strategy and reflection embedded in subject content, not stand-alone learning-to-learn courses. This is why metacognition and self-regulated learning sits as its own competency in the framework. Other competencies rest on varying evidence bases — some strong, some moderate, some thin. Per-competency evidence notes appear under each competency in the curriculum explorer.

For any of these mechanisms to land, four conditions need to be met. The SAFE criteria — Sequenced, Active, Focused, Explicit — are the moderator finding that separates programmes producing real effects from programmes producing near-null effects, confirmed across Durlak et al. (2011) and Cipriano et al. (2023):

  • Sequenced. Skills are built in a deliberate developmental order, with each step assuming and extending the one before.
  • Active. Students engage in active practice, role-play, structured discussion, problem-solving — not passive reception.
  • Focused. Sufficient curricular time is devoted to dispositional skill development as a primary goal, sustained across the school year.
  • Explicit. The specific skills and dispositions are named openly, taught directly, and discussed by name with students.

Meeting these conditions is not automatic. It requires sustained teacher training, and a SAFE-failing implementation produces near-null effects — which is why adoption requires the infrastructure investment named in the limitations section.

Roorda, Koomen, Spilt and Oort’s (2011) meta-analysis of 99 studies and 129,423 students found medium-to-large associations between affective teacher–student relationships and student engagement (r ≈ 0.32–0.39); the 2017 MASEM update confirmed engagement as the mediator. These findings are among the most replicated in the entire literature. The framework’s structural commitment that everyone is a wellbeing teacher, and that relationships are the programme, is grounded directly here.

The teacher as upstream lever

The most empirically vindicated upstream lever is teacher dispositional state and wellbeing. Oberle and Schonert-Reichl’s (2016) study found that ten per cent of variability in students’ morning cortisol occurred at the classroom level, with teacher burnout significantly predicting elevated student cortisol. The CARE for Teachers RCT (Jennings et al., 2017, 2019) produced significant effects on teachers’ emotion regulation, mindfulness, psychological distress and observed Emotional Support, with sustained follow-up effects. Klingbeil and Renshaw’s (2018) meta-analysis of teacher mindfulness-based interventions found Hedges’ g ≈ 0.60. The Jennings and Greenberg (2009) prosocial classroom model — teacher social-emotional competence shapes classroom climate, which shapes student outcomes — is the best-validated theoretical frame for this lever. Yeager’s 2024 mentor-mindset framework explicitly converges with it.

The cumulative direction of this evidence is consistent. The field’s centre of gravity has shifted from student-targeted dispositional curricula to adult relational practice and disciplinary embedding. Yeager’s pivot from student modules to teacher-mediated growth-mindset cultures, the Jennings/CARE evidence on teacher conditions, the Roorda meta-analyses on teacher–student relationships, the Cipriano et al. confirmation that effect sizes shrink as evaluations become independent, the MYRIAD null result for universal adolescent mindfulness, and the EEF’s elevation of metacognition (embedded in subjects) to +8 months progress all point the same way. Dispositional development is not primarily produced by teaching dispositions; it is produced by adults modelling and naming dispositions while scaffolding subject practice in psychologically safe relationships. The curriculum here is designed against that finding rather than around it.

Where this leaves teachers

The framework gives teachers a structure, not a script. Learning targets, KUDs, and criteria are written with enough modularity that teachers can apply them across different projects, performance tasks, and real-world experiences while maintaining a consistent standard for what good performance looks like at each level. Teachers retain the professional judgement to design learning scenarios, decide when students are ready for assessment, and choose the contexts in which performance is observed.

The devil is in the details. Whether teachers apply this framework reliably will vary, and the only way to get good at it is together — through ongoing professional conversations, calibration sessions, and an accumulating bank of exemplars at each level. We are early in this work. Calibration practices are starting; an exemplar bank is starting; inter-rater agreement hasn’t been formally measured yet. We treat this as work in progress and we expect it to take years.

There’s also a less comfortable truth: teachers themselves develop the dispositions they’re trying to support in students. A teacher cultivating self-regulation in students is also working on her own. A teacher who wants students to be honest about their growth has to be willing to be honest about hers. This isn’t separate from teaching the framework — it’s part of what teaching the framework requires.

Reporting

The framework’s design has implications for how reports are structured.

For T1 and T2 learning targets, reporting is at the learning target level, using the five-level descriptors. Each learning target carries its own marker: the level of recent performance, assessed repeatedly across contexts rather than once. The competency-level picture is a profile — a summary of strengths and growth areas across the LTs — not a single aggregated grade.

For T3 dispositional content, reporting happens through Reflection 360 — a structured developmental conversation involving the student, teacher, and (where appropriate) parents, drawing on observation indicators across multiple contexts. The student’s own self-observations are part of this conversation, not as evidence in a tribunal but as material for self-understanding.

We don’t combine dispositional and skill-based data into a single number. Averaging them loses what each is trying to say. A student who is Competent on two T2 learning targets and Emerging on a T3 disposition isn’t best summarised by averaging. The richer report says: this student has substantial competence in the cognitive aspects of this competency and is still developing the dispositional aspects, which is developmentally appropriate at their age and is what we’d expect.

REAL School’s reporting practice is currently evolving toward this design. The school has historically used aggregated grading conventions inherited from prior accreditation contexts; the move to LT-level reporting plus Reflection 360 for dispositions is part of the work this framework is producing. The current state at REAL School is mid-transition.

A note on formative and summative assessment, in Wiliam’s sense: every assessment in this framework is fundamentally formative — it tells the student where they are and what to develop next. Some of those assessments also get reported summatively at term boundaries because reports happen. The same criterion-level descriptor is doing both jobs. The summative use is a use, not a property of the assessment. The goal of summative reporting in this framework is not to rank or to judge. It’s to help students become aware of their own growth and to support their developing agency over their own learning.

Toward agency

All of this — the criteria, the levels, the careful distinctions about how different kinds of learning develop — points somewhere. It points toward agency. A young person who can see their own strengths and growth areas, who has language for what they’re getting better at, and who can set their own goals, is a young person on the path to becoming a self-determined learner. Assessment, done well, is what gets them there: it tells them where they are, what their next step is, and gives them the means to take it.

That last move is where this work points. A young person who can see their own strengths and growth areas, who has language for what they’re getting better at, and who can set their own goals, is a young person on the path to becoming a self-determined learner.

Self-determination at the level of the individual learner only matters if the school as a whole supports the relational ground that makes self-determination meaningful. A self-aware atomised individual is not what flourishing looks like. The framework supports individual development. The school — the relationships, the rituals, the daily texture — supports the community within which individual development is worth doing. Both are needed.

What this approach can’t do

A few honest limitations.

The framework is built on evidence, not yet empirically validated. The design choices are anchored in research on competency development, dispositional learning, and assessment design. Whether the framework as implemented at REAL School produces the outcomes it’s designed to produce is a separate question, requiring measurement over multiple cohorts. We treat this as work in progress.

Inter-rater reliability is real work. Five-level criterion descriptors are only as reliable as the people applying them. The boundaries between adjacent levels are the hardest part — calibration is ongoing, and we are still learning how consistently teachers across the school read the same student work the same way. We treat this as a known measurement challenge, not a solved one.

Dispositions are deeply contextual. Whether a student can self-regulate at school doesn’t tell us whether they can at home, or in conflict with a sibling, or under genuine stress. The situated-cognition tradition (Lave, Wenger and others) makes this point sharply: skill is co-constituted by environment and relationship, not stored neatly inside the learner. The framework tries to handle this through multi-context observation and developmental conversation, but it doesn’t dissolve the limit. Genuinely supporting a young person’s wellbeing requires a tribe — parents, peers, community — that schools alone cannot replicate. We’re honest that the broader social conditions for thriving wellbeing aren’t yet in place; we work with what we have.

The comparison work is Anglophone. The current crosswalk against external frameworks compares REAL School’s wellbeing programme to the UK’s RSHE, the Welsh Curriculum for Wales 2022, and the US-based CASEL framework. The EU’s LifeComp framework is a recognised next addition; Nordic curricula (especially Norwegian and Finnish) and East Asian frameworks (Japanese MEXT) are recognised gaps. This means the framework’s design has been pressure-tested against a particular curricular tradition. Frameworks emerging from very different traditions might surface design tensions that this comparison won’t catch.

Adoption requires substantial school infrastructure. A curriculum is not the same as the conditions that allow it to be taught well. The research evidence we have just laid out has a sharp practical implication that the methodology must name directly: a competency-based wellbeing curriculum makes new demands on teachers, and those demands fall on adults who are themselves under sustained pressure. If the school does not invest in teacher wellbeing alongside the curriculum, the curriculum will not work as designed, and the failure will be misread as a failure of the children rather than of the conditions.

The evidence laid out earlier in this page on teacher dispositional state — Roorda’s meta-analyses, the CARE for Teachers RCT, Yeager’s mentor-mindset framework, Brown et al.’s mentoring meta-analysis — has a sharp practical implication. This places a real obligation on the school. A wellbeing curriculum cannot be adopted as a layer added on top of an already-saturated workload. The work of noticing dispositions, modelling them, naming them in the moment, sustaining warmth and contingency under pressure, and reflecting on practice — the very moves the evidence identifies as the active ingredients — are cognitively and emotionally demanding. They cannot be performed by exhausted, dysregulated, or chronically stressed adults at the level the curriculum requires. If the school’s structural conditions consume teacher capacity rather than support it, the curriculum’s design assumptions will fail to hold.

The implication for adoption is direct. A school taking on this curriculum should expect to invest in teacher wellbeing infrastructure with the same seriousness as in materials and training: protected time for reflection, manageable group sizes, supervision and peer support structures, professional learning that addresses teacher dispositional state rather than only teacher technical skill, and leadership practice that protects teachers from chronic overload. These are not optional climate variables. They are the conditions under which the curriculum’s evidence base actually applies. Without them, what is being implemented is not the curriculum the evidence supports.

We name this here rather than in a footnote because the field has a long history of curricula failing in implementation and the failure being attributed to the wrong cause. The dispositions evidence is clear that universal student programmes deliver small-to-modest effects under good conditions and near-null effects under poor ones. Teacher conditions are a major part of what separates the two cases. Schools considering adoption should plan accordingly, and schools that cannot meet these conditions should not assume the published effect sizes will transfer to their setting. This is also why our work points toward AI as a co-intelligence for teachers — the agentic method page describes the trajectory in more detail.

The models we draw on are not the last word. Broadwell, Haring, and the contemporary cognitive-science literature inform the framework’s design — they don’t determine it. We treat them as useful design frames, and we expect to revise our thinking as understanding improves. Built on contemporary research, designed to be revised as understanding deepens.

This is a work in progress. The framework is more developed in some bands than others, more thoroughly tested in some learning targets than others, and continues to evolve as we use it. Where the research literature is unsettled (especially around dispositional development), we say so on the relevant pages.

References and further reading

Black, P., & Wiliam, D. (1998). Inside the Black Box: Raising Standards Through Classroom Assessment. King’s College London.

Broadwell, M. M. (1969). Teaching for Learning (XVI). The Gospel Guardian.

Brown, A. L., et al. (2023). Meta-analysis of character education programs. Journal of Character Education / Journal of Moral Education, 52(2), 119–138.

Christodoulou, D. (2014). Seven Myths About Education. Routledge.

Cipriano, C., Strambler, M. J., Naples, L. H., et al. (2023). The state of evidence for social and emotional learning: A contemporary meta-analysis. Child Development, 94(5), 1181–1204.

Costa, A. L., & Kallick, B. (2008). Learning and Leading with Habits of Mind. ASCD.

Credé, M., Tynan, M. C., & Harms, P. D. (2017). Much ado about grit. Journal of Personality and Social Psychology, 113(3), 492–511.

Dignath, C., & Büttner, G. (2008). Components of fostering self-regulated learning. Metacognition and Learning, 3, 231–264.

Donker, A. S., de Boer, H., Kostons, D., Dignath van Ewijk, C., & van der Werf, M. P. C. (2014). Effectiveness of learning strategy instruction. Educational Research Review, 11, 1–26.

Durlak, J. A., Weissberg, R. P., Dymnicki, A. B., Taylor, R. D., & Schellinger, K. B. (2011). The impact of enhancing students’ social and emotional learning. Child Development, 82(1), 405–432.

Education Endowment Foundation. (2025). Metacognition and Self-Regulated Learning guidance report.

Flook, L., Goldberg, S. B., Pinger, L., & Davidson, R. J. (2015). Promoting prosocial behavior and self-regulatory skills in preschool children. Developmental Psychology, 51(1), 44–51.

Haring, N. G., Lovitt, T. C., Eaton, M. D., & Hansen, C. L. (1978). The Fourth R: Research in the Classroom. Charles E. Merrill.

Hendrick, C. (2026). Rethinking Retrieval Practice: Remembering Is Not Knowing. The Learning Dispatch.

Henriksen, D., Creely, E., Gruber, N., & Leahy, S. (2025). Generative AI and social-emotional learning: A critical review. Journal of Teacher Education.

Jennings, P. A., Brown, J. L., Frank, J. L., Doyle, S., Oh, Y., Davis, R., et al. (2017). Impacts of the CARE for Teachers program. Journal of Educational Psychology, 109, 1010–1028.

Jennings, P. A., Doyle, S., Oh, Y., Rasheed, D., Frank, J. L., & Brown, J. L. (2019). Long-term impacts of the CARE program. Journal of School Psychology, 76, 186–202.

Jennings, P. A., & Greenberg, M. T. (2009). The prosocial classroom. Review of Educational Research, 79, 491–525.

Klingbeil, D. A., & Renshaw, T. L. (2018). Mindfulness-based interventions for teachers. School Psychology Quarterly, 33(4), 501–511.

Kuyken, W., Ball, S., Crane, C., et al. (2022). Effectiveness and cost-effectiveness of universal school-based mindfulness training. Evidence-Based Mental Health, 25(3), 99–109.

Lave, J., & Wenger, E. (1991). Situated Learning: Legitimate Peripheral Participation. Cambridge University Press.

Lieberman, M. D., Eisenberger, N. I., Crockett, M. J., et al. (2007). Putting feelings into words. Psychological Science, 18(5), 421–428.

Macnamara, B. N., & Burgoyne, A. P. (2023). Do growth mindset interventions impact students’ academic achievement? Psychological Bulletin, 149(3-4), 133–173.

Maynard, B. R., Solis, M. R., Miller, V. L., & Brendel, K. E. (2017). Mindfulness-based interventions for primary and secondary school students. Campbell Systematic Reviews, 13(1), 1–144.

Mischel, W., Ebbesen, E. B., & Zeiss, A. R. (1972). Cognitive and attentional mechanisms in delay of gratification. Journal of Personality and Social Psychology, 21(2), 204–218.

Oberle, E., & Schonert-Reichl, K. A. (2016). Stress contagion in the classroom. Social Science & Medicine, 159, 30–37.

Roffey, S. (2014). Circle Solutions for Student Wellbeing. Sage.

Roorda, D. L., Koomen, H. M. Y., Spilt, J. L., & Oort, F. J. (2011). The influence of affective teacher–student relationships. Review of Educational Research, 81, 493–529.

Roorda, D. L., Jak, S., Zee, M., Oort, F. J., & Koomen, H. M. Y. (2017). Affective teacher–student relationships and engagement: A MASEM approach. School Psychology Review, 46(3), 239–261.

Sadler, D. R. (1987). Specifying and promulgating achievement standards. Oxford Review of Education, 13(2).

Schonert-Reichl, K. A., Oberle, E., Lawlor, M. S., Abbott, D., Thomson, K., Oberlander, T. F., & Diamond, A. (2015). Enhancing cognitive and social-emotional development through MindUP. Developmental Psychology, 51(1), 52–66.

Sperber, J. F., Vandell, D. L., Duncan, G. J., & Watts, T. W. (2024). Delay of gratification and adult outcomes. Child Development, 95(6), 2015–2029.

Tricot, A., & Sweller, J. (2014). Domain-specific knowledge and why teaching generic skills does not work. Educational Psychology Review, 26, 265–283.

Watts, T. W., Duncan, G. J., & Quan, H. (2018). Revisiting the marshmallow test. Psychological Science, 29(7), 1159–1177.

Wigelsworth, M., Lendrum, A., Oldfield, J., Scott, A., ten Bokkel, I., Tate, K., & Emery, C. (2016). The impact of trial stage, developer involvement and international transferability on universal SEL programme outcomes. Cambridge Journal of Education, 46(3), 347–376.

Wiggins, G., & McTighe, J. (2005). Understanding by Design (2nd ed.). ASCD.

Wiliam, D. (2011). Embedded Formative Assessment. Solution Tree Press.

Yeager, D. S., Hanselman, P., Walton, G. M., et al. (2019). A national experiment reveals where a growth mindset improves achievement. Nature, 573, 364–369.

Yeager, D. S. (2024). 10 to 25: The Science of Motivating Young People. Avid Reader Press.

Young, M. (2008). Bringing Knowledge Back In. Routledge.