Every survey you have ever filled out relied on a hidden engineering decision. The moment a form asks you to rate your satisfaction from 1 to 5, or to drag your favourite features into order of preference, the researcher has already chosen a measurement tool that will shape every conclusion drawn from your answer. These tools are called scales, and getting them right is the difference between data that tells a clear story and data that confuses everyone who reads it. Two of the most useful and widely used types are rating scales and rank order scales. Understanding how each works, and when to reach for one over the other, is a core skill for anyone designing a questionnaire.
Table of Contents
What are scales in questionnaires?
A scale is a tool that converts opinions, attitudes, and preferences into numbers that can be measured and compared. Open-ended questions like “What do you think of our library service?” produce rich answers, but they are almost impossible to count or compare across hundreds of respondents. Scales solve this by giving people a fixed set of options that map onto a measurable structure.
The theory behind scaling rests on the idea of levels of measurement. The psychologist Stanley Smith Stevens introduced the now-standard classification in 1946, identifying four levels: nominal, ordinal, interval, and ratio. Each level carries different information and permits different statistical analysis. A nominal scale only names categories, such as the type of device a person uses. An ordinal scale adds order, like a satisfaction ranking from very low to very high, but does not promise equal gaps between points. Interval and ratio scales add measurable, equal distances between values, with ratio scales also having a true zero.
This classification matters because higher-level scales permit more sophisticated statistical analysis than lower-level scales. A good rule is to collect data at the highest practical level, because interval data can always be simplified into categories later, but you cannot upgrade nominal or ordinal data into higher levels after collection.
Why scaling techniques matter
Scales bring consistency and measurability to subjective topics. When everyone answers within the same structure, researchers can calculate averages, frequencies, and distributions that would be impossible with free text. As survey methodologists at Qualtrics explain, rating questions ask respondents to assign a score to each item, while ranking questions ask them to put items in order. Both turn fuzzy human judgement into structured numbers.
There is also a serious risk in choosing the wrong scale. Misapplying measurement levels leads directly to flawed conclusions. One of the most common errors in research is treating ordinal data as if it were interval, assuming the gaps between scale points are equal when they may not be. Knowing your scale type tells you which statistical tests are valid and which will quietly mislead you.
Understanding rating scales
Rating scales ask respondents to evaluate a single item against a fixed set of options. The key feature is that the same rating can be given to many different items. If a respondent loves three different services, they can give all three the top score. This independence makes rating scales fast and comfortable to answer.
Rating scales come in several forms. The numerical rating scale uses numbers, such as rating a service from 1 to 10. The graphic rating scale uses visual markers like stars or smiley faces, which is why customer reviews so often appear as a five-star system. The semantic differential scale places two opposite adjectives at either end of a line, for example “Innovative” to “Uninspired,” and asks respondents to mark where their feeling falls. Each format suits a slightly different purpose, but all share the same logic of scoring items independently.
The Likert scale
The most famous rating scale is the Likert scale. It was introduced by the American social psychologist Rensis Likert in his 1932 work on the measurement of attitudes, and it remains one of the most popular tools in survey research nearly a century later. A Likert item presents a statement, such as “I am satisfied with the customer service I received,” and asks respondents to choose a level of agreement, typically from “Strongly Disagree” to “Strongly Agree.”
The strength of the Likert scale is that it captures the intensity of feeling, not just its direction. Beyond agreement, the same structure can measure frequency, importance, quality, and likelihood, which makes it remarkably flexible. Each response is given a numerical value, allowing researchers to calculate average levels of attitude across a large group.
A long-running debate concerns how many points the scale should have. Research synthesised in the GESIS survey guidelines points to five to seven categories as optimal for reliability, validity, and ease of understanding, since too many categories blur the meaning of each option. A separate literature review on Likert response options found that the five-point scale is the most widely used, with odd-numbered scales of five or seven offering strong reliability and validity. The odd number provides a neutral midpoint, while the choice of length affects how finely you can detect differences.
Limitations to watch for
Rating scales are not flawless. Because the format repeats, respondents can grow tired and select the same answer for every question, a behaviour known as straight-lining. The team at Cint notes that some respondents avoid the extreme options like strongly agree or strongly disagree, gravitating toward the safer middle choices. This central tendency bias, along with social desirability bias where people give answers they think are acceptable, can quietly distort results. Awareness of these effects is essential for interpreting Likert data honestly.
Rank order scales
Rank order scales take a completely different approach. Instead of scoring each item on its own, they ask respondents to arrange a list of items in order of preference or importance. Crucially, each rank can be used only once. If you are ranking five library services, only one can be first, only one can be second, and so on. This forces a genuine trade-off that rating scales do not.
This trade-off is the great advantage of ranking. As Qualtrics points out, because only one item can occupy each rank, the respondent must think carefully about which option truly deserves the top spot. Ranking is therefore ideal when you need to know priorities, not just whether something is liked. A popular and intuitive format is drag and drop, where respondents physically move items into order on screen.
Ranking also resists the satisficing problem that plagues rating scales, since it is impossible to give every item the same score. The result is that ranking scales produce consistent numerical data across all respondents, because everyone is forced to differentiate.
Variations of ranking
Beyond the simple ordered list, researchers use more refined ranking methods. The pairwise method presents two items at a time and asks respondents to choose between them, building a reliable preference order across many comparisons. The MaxDiff method asks respondents to pick the most and least important options from a set, forcing trade-offs that reveal true priorities and avoiding the “everything is important” bias. These techniques are valuable when you have many options but want to keep each individual question simple.
Ranking has its own weaknesses too. It tells you the order of preferences but not the distance between them. A respondent might love their first choice and merely tolerate the rest, yet a rank order alone cannot show that gap. Long lists also become cognitively demanding, since holding ten items in mind and ordering them all is hard work that can lower the quality of answers.
Choosing the right scale
The choice between rating and ranking comes down to what you want to learn. If you want to measure the intensity of an attitude toward each item independently, a rating scale is the right tool. If you want to understand priorities and force respondents to make trade-offs between competing options, a ranking scale serves better.
The practical strengths and weaknesses line up neatly. Rating scales are quick, comfortable, and let people express equal enthusiasm for several items, but they invite straight-lining and central tendency bias. Ranking scales force meaningful differentiation and produce comparable data, but they hide the size of the gaps between choices and tire respondents when the list is long.
You do not always have to pick just one. Researchers often combine the two. As Cint suggests, you can follow a series of rating questions with a ranking question on the same items to uncover the true value behind responses that all received similar scores. This pairing also helps prevent survey fatigue by varying the question format. A common workflow is to use a Likert scale to gauge general agreement, then a ranking question to sort out which items matter most when people are forced to choose.
Matching the scale to your analysis
One final principle ties everything together: choose your scale with the analysis in mind. The level of measurement determines which statistical tests are valid. Nominal data suits frequency counts and the mode, ordinal data supports the median and percentiles, while interval and ratio data permit means, standard deviations, and correlations. Deciding how you will analyse your results before you write the questionnaire prevents the painful discovery, after data collection, that your chosen scale cannot answer your research question. As guidance on scaling techniques advises, you should match the scale to the planned statistical analysis and aim for the highest feasible level of measurement while keeping respondent effort reasonable.
What do you think? If you were designing a questionnaire to find out which facilities students value most in their college library, would you start with rating scales or rank order scales, and why? Could combining both give you a clearer picture than either one alone?
References
- https://statisticsbyjim.com/basics/nominal-ordinal-interval-ratio-scales/
- https://distancelearning.institute/research/educational-research-scaling-techniques/
- https://www.qualtrics.com/blog/rating-or-ranking-choosing-the-best-question-type-for-your-data/
- https://journalism.university/communication-research-methods/scales-of-measurement-nominal-ratio/
- https://www.simplypsychology.org/likert-scale.html
- https://www.gesis.org/fileadmin/admin/Dateikatalog/pdf/guidelines/design_rating_scales_questionnaires_menold_bogner_2016.pdf
- https://files.eric.ed.gov/fulltext/EJ1369114.pdf
- https://www.cint.com/blog/using-rating-questions-vs-ranking-questions-in-a-survey/
- https://www.surveyking.com/blog/rating-scale/

Leave a Reply