Numbers feel comfortable. Height, weight, income, marks in an exam – these slot neatly into formulas and statistical tables. But a huge part of social science research deals with things that refuse to behave like numbers: a person’s loyalty to a brand, a student’s anxiety about exams, a community’s trust in local government. These are qualitative variables, and measuring them is one of the genuine puzzles of research methodology. This post breaks down how researchers actually capture such non-numeric characteristics, why validity and reliability sit at the centre of the whole exercise, and how attitude scales – especially the Likert scale – turn opinions into data you can analyse.
Table of Contents
- Why qualitative data resists straightforward measurement
- The risk of researcher bias
- Defining the variable first
- Validity and reliability: the two tests every tool must pass
- Validity: are you measuring the right thing?
- Reliability: are your results consistent?
- Attitude scales: turning opinions into numbers
- The Likert scale
- The Thurstone scale
- The Guttman scale
- The semantic differential scale
- Where these scales are applied
- A caution: response bias
- Choosing the right technique
Why qualitative data resists straightforward measurement
Qualitative data is descriptive and meaning-focused. It deals with the “why” and “how” behind behaviour rather than the “how many.” Variables like social attitudes, cultural values, motivation, and group behaviour do not come with a built-in unit of measurement. There is no kilogram of patriotism and no metre of job satisfaction.
This creates a real problem. A researcher studying voter trust or workplace morale still needs evidence that can be compared across people and analysed systematically. Pure description alone – pages of interview notes – is rich but hard to summarise. So the central challenge becomes: how do you assign some kind of structure or number to a characteristic that is, by nature, subjective?
The risk of researcher bias
Because qualitative data depends on human experiences and interpretation, there is always a chance that the researcher’s own bias shapes how data is collected or read. Two researchers can interpret the same interview differently. Qualitative measurement is partly an attempt to reduce this subjectivity by introducing systematic procedures – coding schemes, standardised questions, and rating scales – that anyone following the same rules would apply in roughly the same way.
Defining the variable first
Before any measurement tool is chosen, the abstract concept must be defined clearly. If a study aims to measure “empowerment,” that word means nothing until the researcher specifies what it looks like in their context – financial independence, decision-making power, or freedom of movement. This step, called operationalisation, converts a fuzzy idea into something observable. Only after defining the variable can you decide on a technique to measure it.
Validity and reliability: the two tests every tool must pass
No matter how clever a measurement method looks, it is useless unless it satisfies two conditions. These two ideas are the backbone of measurement in research.
Validity: are you measuring the right thing?
Validity asks whether the tool actually captures the characteristic it claims to capture. A questionnaire meant to measure exam anxiety should measure anxiety – not general intelligence or study habits. For attitude scales, validity is often checked by consulting a panel of experts who judge whether the statements truly reflect the concept under study. A tool can produce consistent results and still be invalid if it is consistently measuring the wrong thing.
Reliability: are your results consistent?
Reliability refers to consistency. If the same group answers the same scale under similar conditions, the results should be stable rather than random. A common way to check internal reliability is by calculating Cronbach’s alpha coefficient, which indicates how well the different items on a scale hang together as a measure of one underlying concept. Building reliable and valid items requires careful attention to wording and structure along with proper psychometric testing.
Attitude scales: turning opinions into numbers
The most widely used solution to measuring qualitative variables is the attitude scale. Attitude measurement is the systematic assessment of how people evaluate objects, ideas, people, or events. The core trick is scaling – assigning numbers to degrees of opinion so that the data can be statistically analysed. Several established procedures exist, and each handles the problem slightly differently.
The Likert scale
The Likert scale is by far the most popular attitude-measurement tool. It was introduced by Rensis Likert in 1932 and remains a standard across the social sciences nearly a century later. Respondents are presented with a statement and asked how strongly they agree or disagree, usually on a five-point continuum: strongly disagree, disagree, neutral, agree, strongly agree.
Each response is assigned a number – for example, 1 to 5 – so that the numerical values can be used to measure the attitude under investigation. By summing or averaging scores across several related statements, the researcher arrives at an overall measure of a person’s attitude. This is why it is sometimes called a “summated” scale. Beyond simple agreement, Likert-type items can also measure frequency, importance, quality, and likelihood, which makes them flexible.
A practical design point: most Likert scales use an odd number of response options so that a neutral midpoint exists. Research generally favours five- to seven-point scales, and evidence suggests that odd-numbered scales of more than five points, especially seven points, perform well on reliability and validity. Fully labelling each point also tends to improve reliability.
The Thurstone scale
The Thurstone scale, created by Louis Leon Thurstone in 1928, was actually the first formal technique developed to measure attitudes. Building it is elaborate. Researchers generate a large pool of statements, and a panel of judges rates each one for how favourable or unfavourable it is, typically on an eleven-point range. Statements the judges disagree on are dropped, and the survivors are chosen so they sit at roughly equal intervals along the attitude continuum. Respondents then simply tick the statements they agree with, and their score is based on the scale values of those items.
The method offers precision, but it is costly and time-consuming to construct. In fact, the Likert scale was developed partly to simplify the heavier mathematics of the Thurstone approach, which is one reason Likert overtook it in everyday use.
The Guttman scale
The Guttman scale, also called the cumulative scale, was developed by Louis Guttman in the 1940s. Its logic is hierarchical: statements are ranked from mild to extreme, so that agreeing with a strong statement implies agreement with all the weaker ones before it. If a respondent agrees with the eighth statement in a list, it suggests they also agree with statements one through seven and disagree with the more extreme ones that follow. Its strength is that it reveals whether an attitude forms a clean, cumulative pattern, though real-world data rarely fits this model perfectly.
The semantic differential scale
The semantic differential, developed by Charles Osgood and colleagues, takes a different angle. Instead of agreement statements, respondents rate a concept on a series of bipolar adjective pairs – for example, good-bad, weak-strong, friendly-unfriendly – usually across a seven-point spread. This scale is especially good at capturing the emotions and connotations associated with a concept, which makes it a favourite in brand-perception and image studies.
Where these scales are applied
Attitude scales are central to disciplines where measuring opinion is the whole point. In psychology, they assess constructs like anxiety, self-esteem, and motivation. In education, they capture students’ attitudes toward subjects, teaching methods, or institutions – research on attitudes toward mathematics, for instance, leans heavily on Likert-type instruments. In sociology, they measure attitudes toward social issues such as religion, gender roles, and community trust.
The reach extends well beyond academia. Market researchers use these scales to gauge brand perception, healthcare studies use them to record self-reported pain or anxiety, and organisations use them in employee-satisfaction surveys. The common thread is the conversion of a subjective, qualitative variable into structured data that supports comparison and statistical analysis.
A caution: response bias
These tools are powerful but not flawless. Self-reports of attitudes are sensitive to small changes in question wording, format, and order. A well-known threat is social desirability bias, where people answer in ways that make them look good rather than telling the truth – few respondents would openly admit to a prejudice, for example. Offering anonymity on self-administered questionnaires helps reduce this pressure. Another issue is acquiescence bias, the tendency to agree regardless of content, which researchers counter by mixing positively and negatively worded statements. Recognising these biases is part of designing a measurement tool that is genuinely valid.
Choosing the right technique
There is no single best scale. The choice depends on the research question, available resources, the population, and the nature of the attitude being studied. Likert scales offer simplicity and broad applicability, which makes them suitable for most general purposes. Thurstone scales provide more precision but demand far more development effort. Guttman scales work best when an attitude naturally forms a cumulative hierarchy. Semantic differentials shine when the goal is to capture emotional and evaluative dimensions. A thoughtful researcher matches the tool to the job rather than defaulting to the most familiar option.
What do you think? If you were designing a study on students’ attitudes toward online learning, which scale would you trust most – and how would you guard against respondents simply telling you what they think you want to hear? Can a subjective feeling ever be measured with the same confidence as height or weight, or does some richness always get lost in the numbers?
References
- https://blog.upmetrics.com/essential-qualitative-data-collection-methods
- https://blog.upmetrics.com/qualitative-measurement
- https://journals.plos.org/plosone/article?id=10.1371%2Fjournal.pone.0239626
- https://pmc.ncbi.nlm.nih.gov/articles/PMC7529215/
- https://www.simplypsychology.org/likert-scale.html
- https://www.cogn-iq.org/learn/theory/attitude-measurement/
- https://www.ijem.com/number-of-response-options-reliability-validity-and-potential-bias-in-the-use-of-the-likert-scale-education-and-social-science-research-a-literature-review
- https://bns.institute/behavioural-sciences/measuring-attitudes-methods-scales/
- https://theintactone.com/2019/02/19/rm-u3-topic-5-thurstone-scale-likert-scale-and-semantic-differential-scale/
- https://www.leadquizzes.com/blog/semantic-differential-scale/
- https://www.questionpro.com/blog/thurstone-vs-guttman-scale/
- https://dovetail.com/surveys/semantic-differential-scale/
- https://www.researchgate.net/publication/276394797_Likert_Scale_Explored_and_Explained

Leave a Reply