Working Paper 00214 min read

Measurement Without Diagnosis

Ethical and regulatory boundaries in commercial behavioural assessment

Brett DoddsInstitute of Behavioural Performance

Abstract

Commercial behavioural assessment operates in a space adjacent to clinical psychology without being equivalent to it and without always being subject to the same professional regulation. This paper sets out where we understand the boundary to lie, why we treat it as binding even where it is weakly enforced, and what operational constraints follow for instrument design, report generation, data handling and practitioner licensing.

The boundary is not primarily a legal question about title protection. It is an evidentiary question about what an instrument’s validity evidence supports. An instrument developed to describe behavioural tendency does not acquire diagnostic authority because it is administered to someone who happens to be unwell.

We derive seven operating constraints from this position. Four are implemented in the current Behavioural Performance Index. Three are commitments that are incomplete or not yet built. We also examine the hazards introduced by machine-generated report prose and set out where our own practice currently falls short.

The paper is offered partly as a statement of our own practice and partly as a proposed baseline for a category that currently has no consistent one. We have tried to be accurate about which of the two any given claim belongs to.

Keywords

  • test-use ethics
  • scope of practice
  • commercial psychometrics
  • construct validity
  • AHPRA
  • ITC guidelines
  • generative AI

1. Why this paper exists

The commercial behavioural assessment market sells instruments that describe personality, motivation and working style to people who may also be under significant psychological strain. Business owners are a substantial segment of that market. Some studies report high levels of exhaustion, poor recovery and distress in founder populations, alongside barriers to professional help-seeking including stigma. The size of that problem varies by study and population, and the evidence base remains uneven. The risk does not depend on an exact prevalence figure.

Combine three things. A person under strain, a reluctance to seek clinical help, and an available product that produces authoritative-sounding psychological descriptions of them. The hazard is specific. An assessment report can function as a substitute for help-seeking. It can provide a satisfying explanation for experiences that warrant attention beyond the assessment. It can do this without any intent on the publisher’s part, purely as a consequence of the format and the authority the reader assigns to it.

The category has an obligation to address this explicitly rather than rely on a footer disclaimer. This paper sets out our position.


2. The regulatory landscape, accurately stated

It is worth being precise about what is and is not regulated, because vagueness here serves commercial interests and we would rather not benefit from it.

2.1 Title protection in Australia

In Australia, the title psychologist is protected and practising as a psychologist requires registration with the Psychology Board of Australia, supported by the Australian Health Practitioner Regulation Agency. The Board’s stated priority is public safety and its professional competencies include psychological assessment as a core area of practice.

That framework governs registered psychologists and the use of protected titles. Commercial behavioural questionnaires published outside clinical practice do not automatically become psychological services merely because they ask about behaviour. That does not leave them unrestrained. Consumer law, privacy law, representations about expertise, intended use and the nature of the information collected may all remain relevant.

The Institute of Behavioural Performance is not a clinical service, does not represent the BPI as a clinical instrument and does not claim diagnostic authority.

This section states our operational understanding rather than legal advice. The final wording and the product architecture should be reviewed as the instrument expands into new jurisdictions or use cases.

2.2 The guidance that applies as a standard of practice

The International Test Commission’s guidelines represent international consensus on responsible test development and use. They are framed in terms of the competencies test users need and the standards publishers should meet. The ITC does not act as a national regulator, but its guidance covers test use, adaptation and technology-based assessment, including validity, fairness, security and privacy.

The Standards for Educational and Psychological Testing, jointly issued by AERA, APA and NCME, remain a central reference for validity evidence and appropriate score interpretation.

These documents may not bind every commercial publisher as law. We treat their central principles as binding anyway, for the reason set out next.

The temptation for a commercial publisher is to reason as follows: we are not regulated in the same way as psychologists, therefore the boundary is wherever our lawyers say it is. That is a category error.

The real constraint is evidentiary. An instrument’s claims are licensed by its evidence and by nothing else. If the BPI is supported as a description of behavioural tendency and current self-reported capacity in business owners, then that is the full extent of what it may claim.

Not because a regulator would necessarily object, but because claims beyond the evidence are unsupported. Diagnostic authority is not conferred by disclaimer, professional tone, practitioner licensing or market position. It requires evidence for diagnostic use. We do not have that evidence and are not seeking that use.

This reframing produces a boundary that does not disappear when the legal environment is permissive. It also travels more cleanly across jurisdictions.


3. Seven operating constraints

From the position above we derive seven constraints governing instrument content, scoring, report generation, data handling and practitioner licensing. Four are implemented in the current product. Three are incomplete or not yet built, and we mark them as such rather than describe them in the present tense.

3.1 Implemented

C1 — No clinical constructs in the instrument. The BPI measures 25 behavioural traits across five domains, none of which is intended as a clinical construct. It does not include scales for depression, anxiety, trauma, substance use or suicidality, and contains no DSM- or ICD-referenced construct. The nearest-adjacent traits are Emotional Regulation, Resilience Reset, Optimism and Coachability, all defined operationally: steadiness under pressure, recovery after setbacks, expectation that progress is possible and willingness to adjust in response to feedback.

We note a limitation. Construct-level cleanliness does not guarantee item-level cleanliness. An item can function like a mood or distress screen regardless of the trait it sits under, and the four traits named above are where that drift is most likely. A systematic item-level review against this criterion remains outstanding and is the first thing we would want an external reviewer to check.

C2 — No diagnostic, prognostic or clinical-adjacent language in output. Reports avoid terms implying pathology, disorder, syndrome or clinical severity, and do not characterise a respondent as impaired or dysfunctional. The energy sections report energy associated with categories of work, not burnout. Burnout is an established occupational-health construct with specialist measures. The BPI does not purport to assess it.

The framing is deliberately anti-pathologising. The lowest score band is described as a system and support priority rather than a personal flaw. Where a pattern appears to be costing the business, the report states that more discipline may not solve it. Low scores are routed toward structure, role design, ownership and support rather than personal deficiency.

C3 — No causal claims about mental state. Reports describe measured patterns and possible operational consequences. They do not assert why a respondent is in a particular state and do not attribute that state to psychological history, relationships, health or personal characteristics beyond the measured constructs.

The report also declines quantified consequence claims. It identifies where to investigate. It does not claim a precise financial loss, clinical outcome or forecast unless separate evidence supports that use.

C4 — Explicit scope statement in the reader’s path. The report carries a positioning note within the body, not only in a footer or terms document. It states that the BPI is a commercial self-assessment and owner-operating report, not a psychological, clinical, hiring, medical, financial or legal diagnostic tool. Results are decision-support signals to be considered alongside the reader’s judgement, business data and, where appropriate, professional advice. The appendix repeats the limitation.

3.2 Incomplete or not yet implemented

C5 — Welfare signposting on defined triggers. Not currently built. The intended response is a short, calm welfare block when scoring indicates persistently low energy across multiple domains or other pre-defined patterns that merit additional attention.

The wording must not imply that the instrument has detected illness, severity or clinical risk. It should state that the assessment cannot determine the cause of the pattern and suggest speaking with a qualified health professional where it is sustained, worsening or affecting daily functioning.

The trigger should be rule-based rather than left to the discretion of a report generator or practitioner. The threshold, wording and escalation logic require review by appropriately qualified clinical and legal professionals before implementation.

We regard this as the most significant gap in our current practice. An owner whose energy scores are at the floor across every domain currently receives role-design guidance and nothing else. That guidance may be commercially appropriate and still insufficient as welfare signposting.

C6 — Licensed practitioners inherit the boundary. Not yet applicable. Practitioner licensing is not live. When it is, practitioners will be contractually bound to the same interpretive limits as the instrument. Licensing will not confer latitude beyond the evidence base. A coaching qualification does not license clinical interpretation of a non-clinical instrument.

Contractual wording alone will not be enough. Training, supervision, complaint handling and removal of licence will need to make the boundary operational.

C7 — Privacy and data governance match the sensitivity of the information. Partly implemented and not yet independently audited. Behavioural assessment data may reveal information a respondent considers highly personal even where the instrument is non-clinical. The product should therefore minimise collection, state the purpose of processing, restrict reuse, define retention and deletion periods, protect identifiable data and disclose any third-party or overseas processing.

Where language models or other external services are used, identifiable information should be excluded wherever possible. Assessment responses should not be used to train external models unless the respondent has given specific and informed consent.


4. The machine-generated prose problem

This section addresses a hazard specific to contemporary assessment products and only partly addressed by standards developed before widespread generative AI.

Where narrative report text is generated by a language model rather than assembled from a fixed library of approved statements, three failure modes arise. All three are boundary violations even when no individual sentence looks obviously clinical.

The first is extrapolation beyond the evidence. A generative model asked to write insightfully about a score profile will produce claims that sound like interpretation but are not licensed by the scoring. It may infer motive, history or internal experience because those make for better prose. Each such inference is an unsupported claim delivered in the instrument’s voice.

The second is clinical register drift. Models trained on general text have absorbed clinical and therapeutic language. Without tight constraint, generated prose can drift toward therapeutic framing, pathologising metaphor and the interpretive posture of clinical formulation even while avoiding formal diagnostic terms.

The third is authority laundering, and it is the most serious because it is difficult for the reader to see. Prose generated from a supported score inherits the perceived authority of that score whether or not the specific narrative claim is supported. The reader cannot easily distinguish deterministic scoring from generated commentary.

Our current response is architectural. Scoring is fully deterministic and contains no generative component. Narrative generation is separated from scoring and constrained by scope, prohibited inference classes, structural validators and banned-language checks before output enters the report.

That separation matters, but it does not solve the problem by itself. Description is part of interpretation. A deterministic score followed by unsupported commentary can still mislead the reader.

A more defensible architecture therefore requires an approved claim set tied to constructs and score bands, explicit limits on profile-combination inferences, adversarial testing, version control, logging and a safe fallback where the narrative fails validation. Machine-generated wording should also be disclosed to the reader.

Our current system moves in that direction but should not be presented as proven to prevent unsupported inference. It is our current view of a defensible minimum, not an established category standard.


5. Intended use and prohibited use

Scope statements are stronger when they specify what the instrument is for and what it is not for.

The BPI is intended to support:

  • owner self-reflection;
  • discussion of role design, systems and support;
  • identification of areas requiring operational investigation;
  • coaching within the stated construct boundaries;
  • reassessment of current energy and work demands.

It is not intended to support:

  • diagnosis or treatment planning;
  • recruitment screening or employee selection;
  • promotion, termination or disciplinary decisions;
  • fitness-for-work determinations;
  • emergency or suicide-risk assessment;
  • insurance, legal-capacity or medical decisions;
  • claims about whether someone is a good or bad business owner.

These limits should appear in product documentation and practitioner agreements, not only in this paper.


6. Where the boundary genuinely blurs

Three cases resist clean resolution. We state them rather than pretend the framework handles everything.

Sustained high load is the hardest, and it is the case we currently handle worst. A respondent reporting persistently low energy across domains may be experiencing something clinically significant. The instrument cannot determine that and it would be a boundary violation to try. But saying nothing where a concerning pattern is plainly visible is also a failure.

C5 is the intended answer: surface a general welfare prompt and make no clinical claim. Until it is built, the report treats a floor-level profile the same way it treats any other profile. That is a gap we are naming rather than defending.

Self-diagnosis by respondents is the second. Readers routinely map non-clinical descriptions onto clinical categories encountered elsewhere. We cannot prevent this. We can avoid language that invites it and state scope clearly, but the residual risk is real.

Practitioner drift is the third. A licensed practitioner with a coaching background, working with a distressed client over multiple sessions, faces continuous pressure toward clinical interpretation. Contractual constraint addresses this weakly on its own. Practitioner training, supervision and quality assurance remain unresolved parts of our proposed licensing model.


7. A proposed category baseline

We put forward seven commitments as a minimum standard for commercial behavioural assessment. We do not claim originality for them. We would prefer competitors met them too, and we recognise that publishing them creates a standard against which our own practice can be checked.

  1. Publish the construct list. State what is measured. Instruments that do not disclose their constructs cannot be evaluated for scope.
  2. Publish the validity evidence, including its absence. Where evidence is pending, say so. A pre-registered protocol is stronger disclosure than a vague claim of validation.
  3. State intended and prohibited uses. Scope should be operational, not implied.
  4. Put the scope statement in the reader’s path. Not only in terms of service.
  5. Disclose and constrain generative components. State where machine-generated text is used and how unsupported claims are prevented.
  6. Operate a welfare-signposting protocol. Triggered by defined rules and reviewed by appropriately qualified professionals. We do not currently meet this commitment.
  7. Publish data-governance principles. State what is collected, why it is collected, how long it is kept, where it is processed and whether it is reused.

Practitioner licensing, where offered, should inherit every one of these commitments rather than operate as a separate interpretive layer.


8. Limitations

This paper describes a position and a set of practices. It does not report evidence that those practices are effective.

We do not know empirically whether our scope statements are read or understood. We have not yet shown that our validator machinery prevents clinical register drift or unsupported inference across the full range of score profiles. A welfare prompt has not been built, so its effects on help-seeking, distress or false reassurance are unknown.

We have also not completed an item-level review of the instrument against C1. The construct list is designed to be non-clinical. The 175 items have not yet been systematically audited for clinical drift.

Our privacy and data-governance controls have not been independently audited against the full set of commitments proposed here.

We are also an interested party proposing a standard that reflects choices we have already made. That is a real conflict. The mitigation on offer is specificity: the commitments in Section 7 are concrete enough to be checked, including against us.


References

American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for Educational and Psychological Testing. Washington, DC: AERA.

Bartram, D., & Hambleton, R. K. (2016). The ITC guidelines: International standards and guidelines relating to tests and testing. In F. T. L. Leong et al. (Eds.), The ITC International Handbook of Testing and Assessment. Oxford University Press.

International Test Commission. (2001). International guidelines for test use. International Journal of Testing, 1(2), 93–114.

International Test Commission. (2006). ITC Guidelines on Computer-Based and Internet-Delivered Testing.

International Test Commission. (2017). ITC Guidelines for Translating and Adapting Tests (2nd ed.).

International Test Commission & Association of Test Publishers. (2025). Guidelines for Technology-Based Assessment.

Psychology Board of Australia. (2025). Professional Competencies for Psychologists. Australian Health Practitioner Regulation Agency.

Psychology Board of Australia. Registration Standards and Guidelines. Australian Health Practitioner Regulation Agency.

Shum, D., Kinsella, G., et al. (2025). Psychological testing in the profession of psychology: An Australian study. Australian Psychologist.


Institute of Behavioural Performance working papers are pre-publication documents circulated for comment. Correspondence: behaviouralperformance.org