Saludos Psychology Group

Dr. Kimberly Fitzgerald González

Licensed Clinical Psychologist

Miami - Los Angeles

FL PY10967 - CA PSY31536

Clinical Considerations

Test Construction and Validity

A psychological test is a scientific instrument. What it can tell you about a person depends entirely on how it was built, who it was built on, and whether the question being asked of it is one it was designed to answer.

Why construction matters at the point of use

Every administration accepts a set of assumptions

Administering a psychological test means accepting the assumptions built into it: what it measures, who it was designed for, what a given score signifies, and the conditions under which results are interpretable.

Most clinicians use instruments they did not build, which is normal. Using one well requires knowing how it was built — what construct it targets, how items were selected, who comprised the normative sample, what the validity evidence actually supports, and where the instrument's appropriate use ends.

That knowledge is what separates a score from a finding.


The construct and the items

Where most testing error originates

Every test begins with a construct — an abstract concept that cannot be directly observed. Intelligence. Depression. Anxiety. None can be placed under a microscope. What can be done is to define the concept precisely enough to measure its behavioral and experiential expressions, then build items intended to capture them systematically.

The items are the developer's attempt to bridge construct and observable behavior, and the bridge is never perfect. A depression scale measures what its developers understood depression to look like, in the population available to them, at the time of writing. Whether it captures depression as it presents in a particular patient, from a particular background, at a particular point in life, is a separate question.

Factor analysis is the tool that evaluates the fit. Given a pool of candidate items, it reveals whether they cluster into coherent underlying dimensions or are quietly measuring several things at once. A well-constructed depression scale may resolve into distinct factors — cognitive, somatic, anhedonic. Where items fail to cluster cleanly, either the construct is poorly specified or the items are doing something unintended.

Confirmatory factor analysis extends this, testing whether a proposed structure holds when the instrument is administered to a new population. An instrument whose factor structure replicates across diverse samples has internal architecture that holds. One whose structure collapses in a new sample is reporting its own limits — and indicating when it should not be used.


Validity

A body of evidence, not a property

Validity is the degree to which an instrument measures what it claims to. It is not a single attribute a test possesses or lacks. It accumulates across studies, populations, and use contexts, and it is always specific: valid for particular purposes, with particular populations, under particular conditions.

Types of validity evidence

  • Content validity — whether items represent the full range of the construct; a depression measure capturing sadness while omitting cognitive, somatic, and anhedonic features has a content problem
  • Construct validity — whether the instrument behaves as theory predicts, correlating with what it should and not with what it should not; the foundation of modern validity theory, following Cronbach and Meehl
  • Criterion validity — whether it predicts what it purports to predict; a risk instrument should predict the outcome it was built for, and where it does not, the claim is weak
  • Cultural validity — whether the same construct is measured in the same way across groups; the question least often asked and most consequential for diverse populations

Reliability

Necessary, and not sufficient

Reliability concerns consistency: whether an instrument produces the same result under the same conditions, whether raters score it equivalently, whether results remain stable when the underlying construct has not changed.

It is a precondition for validity — accuracy is impossible without consistency. But an instrument can be highly reliable and measure the wrong thing with great consistency.

A reliable instrument administered to the wrong population, or used for an unvalidated purpose, produces stable numbers and errs in the same direction every time.


The normative sample

The population standing behind every score

Every score is a comparison. An 85th percentile result means the person scored above 85% of the standardization sample. That sample is present in every interpretation, whether or not it is examined.

Where the sample does not resemble the patient — demographically, linguistically, culturally, clinically — the comparison degrades. An instrument normed decades ago on a narrow population is not the same instrument when administered outside that population, even though the scoring procedure is identical.

Questions any normative sample should answer

  • When were the data collected, and has the construct shifted since?
  • What was the composition of the sample, and does it correspond to this patient?
  • Were clinical populations included, or community samples only?
  • Has the instrument been separately validated with culturally diverse populations?
  • Do subgroup norms exist that are relevant here?

How this is handled here

Instrument selection as a clinical decision

Instrument selection at this practice is treated as a decision requiring justification rather than a default. Which measure, whether its normative data suit this patient, what the validity evidence actually supports for this purpose, and where the instrument's limits require clinical judgment to fill the gap.

Where assessment results carry consequences beyond the consulting room — in legal proceedings, in placement decisions, in documentation that follows a person — the appropriateness of the instrument to the person is part of what makes the finding defensible.

A test score is a starting point. What it means depends on what the instrument was built to measure, who it was built on, and whether the interpretation accounts for the distance between the two.

New patients are seen by appointment. No referral required.

Schedule an Appointment →

This page is for educational purposes only and does not constitute clinical advice, diagnosis, or treatment. If you are experiencing a mental health crisis, call or text 988 to reach the Suicide and Crisis Lifeline. In a medical emergency, call 911 or go to the nearest emergency room.