Beyond “Doing Better”: Improving the Objectivity of Cat Behavior Assessment

Issue 31 | July 2025

Written by Jacklyn J. Ellis, PhD, CAAB, CSB-C


Abstract

At Toronto Humane Society, inconsistencies in how cat behavior and welfare were being reported led to the development of a new system using four simple, standardized behavior rating scales. These ordinal rating scales — measuring fear, anxiety, and stress; response to petting; participation in play; and food intake — use a 0 to 5 scale to track subtle changes over time. This approach allows staff and volunteers to report behavior more clearly and consistently, making it easier to monitor progress and make informed decisions about care, interventions, and placement.

The scales were designed to be easy to use in a busy shelter environment while still providing meaningful data. A key part of the system’s success is training — ensuring that different people interpret behaviors the same way. The author has recently released a free online version of the training offering CEU credits from IAABC and CCPDT, now available to anyone interested in applying the scales in shelters, clinics, or homes.

Since implementing this system, Toronto Humane Society has seen more efficient case management and increased feline welfare, ultimately improving adoptions. The scales have helped guide adjustments to behavior plans, evaluate the efficacy psychopharmaceuticals, and support decisions for alternative placements. By offering a reliable way to monitor feline welfare, this tool can help shelters and professionals everywhere better understand and support the cats in their care.

The Need for Consistent and Discrete Data to Assess Cat Behaviour and Welfare

When I first started at Toronto Humane Society and set about evaluating our existing methods for monitoring the behavior and welfare of the cats in our shelter, I found many wonderful things. We had a large number of dedicated volunteers who spent countless hours working with the cats in our care. And these volunteers wrote intricately detailed notes about how each interaction went. But they weren’t always reporting the information I needed.

I found I was getting phrases like “she’s doing better,” “he let me pet him,” or “I wiggled the wand toy and gave her treats,” but that wasn’t clear enough for me — What behavior was getting better? And by how much? Did the cat enjoy the petting or merely tolerate it? Did she engage with the wand toy and/or consume the treats?” But I couldn’t exactly blame the volunteers, since I hadn’t actually told them what information I wanted, or how I wanted them to convey it to me.

That’s when I decided to create this program. I focused on four key behavioral indicators of well-being, and asked volunteers to report on those behaviors using a 0 to 5 rating scale, where 0 was the absolute best and 5 was the absolute worst. The levels of each scale were operationally defined, and a training system was created to ensure reliability between raters — meaning that if one of my volunteers scored a cat as a 3 on one of the scales, I could be reasonably certain that I would have scored the observed behaviours a 3 for that scale as well. This system has completely revolutionized the way we monitor feline behavior and welfare at Toronto Humane Society, allowing my case management to be more effective and efficient, and ultimately getting cats into adoptive homes much more quickly.

I first published my findings in an article titled “Beyond ‘Doing Better’: Ordinal Rating Scales to Monitor Behavioural Indicators of Well-Being in Cats,” in the journal Animals in 2022, and I am so excited to share some developments that have taken place since then with the readers of IAABC Foundation Journal.

The Four Ordinal Rating Scales: Measuring Behavior and Welfare

Ordinal data is a type of categorical data where the possible responses are ranked in a meaningful order, but the distances between the categories are not necessarily equal or known. I thought ordinal data was well suited to assess these key behavioral indicators of well-being because these scales would offer more sensitivity to subtle changes in a behaviour than simple binary presence/absence data. For example, how active in a day are you on a scale of 1 to 10 offers a more detailed understanding of your daily activity level than asking “Are you active in a day? Choose yes or no.” The use of these scales would also be more feasible in a shelter setting because they would require fewer resources (e.g., person-hours or technology) than the continuous quantitative behavioural observations via video recordings often collected to get a thorough understanding of a behaviour in academic research (e.g., measuring your activity level by watching a 24-hour video of you and coding the duration, frequency, and intensity of your activity bouts).

The first scale was the Fear, Anxiety, and Stress (FAS) score. Now, of course, I didn’t actually develop this scale, but rather adapted it from the Fear Free Pets program (Fear Free, n.d.). This scale is intended to reflect the more traditional body language that a cat uses to communicate with us how fearful, anxious, or stressed they are feeling. The Spectrum of Feline Fear, Anxiety, and Stress infographic available from the Fear Free Pets website describes certain behaviours as a range of numbers, but for our purposes we needed a little more precision, so I made a few slight modifications.

Next, I developed the Response to Petting (RTP) score and the Participation in Play (PIP) score. These two behaviors — play and petting — often increase or decrease alongside behaviours typical of fear, anxiety, and stress, but not always. I found that if we only monitored the FAS score, we were missing valuable nuances in behavioral changes specifically related to petting and play. Additionally, having more information about how much a cat enjoys petting or play allows us to evaluate if these activities can be used as reinforcers later on.

The fourth scale, Food Intake Summary (FIS) score, allowed us to track food consumption as an indicator of stress and well-being. A decrease in food intake can be an early sign of stress or illness, making this scale a critical tool for tracking changes over time. Of course, our shelter was already tracking food intake, but I found we were collecting so much data on food intake that it was difficult to get a clear picture of how well the cats were eating at a glance. Transforming the complicated data we were already collecting to be reflected in the same scale we were using to monitor the other key indicators allowed a holistic interpretation of all four together in the same graph.

These scales were designed to be clear, efficient, and feasible for shelter staff, volunteers, and other animal welfare professionals. They provide a standardized approach to evaluating progress, making it easier to determine the success of intervention strategies or make (and potentially justify) decisions related to feline placement.

Inter-observer Reliability and Agreement

Of course, none of these scores would mean anything if people used them inconsistently. If my 3 on the RTP score isn’t the same as your 3, then the data becomes unreliable.

Inter-observer reliability is essential in research, especially when employing observational methods, as it ensures consistency in data recording across multiple observers. There is a saying in statistical analysis: garbage in, garbage out. Essentially, this means that even the most sophisticated data analysis methods will yield meaningless results if the data collected is not reliable. Ensuring inter-observer reliability is crucial for maintaining the validity and credibility of study results by reducing bias and preventing findings from relying solely on one observer’s interpretation. Ultimately, it verifies that different researchers assess the same phenomenon in a uniform manner.

To assess whether my scales were designed with enough specificity that they could achieve acceptable inter-observer agreement and reliability, I trained 16 raters on each scale and then had them each watch 30 short video clips, scoring them using the scales from 0 to 5. I then compared their scores against my own, which served as the reference standard for agreement and reliability.

Amazingly, we found that the average inter-observer agreement for each scale was almost perfect, and the average inter-observer reliability was excellent. These are the kind of results a researcher dreams of, and gives me reasonable certainty that I can trust that a volunteer’s score is a meaningful and accurate indicator of the cat’s behavior on that day.

This level of reliability is critical in any setting where data is being used to monitor the success or failure of interventions, make pathway decisions (such as barn placement), and track the progress of individual animals over time. By using these scales, I can minimize the subjectivity and variability inherent in the way volunteers used to report how a visit went, leading to more informed, evidence-based decisions. That being said, the volunteers still write their free text observations of the interaction. This information can be used to aid in interpreting the scores (perhaps a dip in scores can be explained by a note in the volunteer comments indicating there was construction in the room next door that day) and to help us craft adoption bios with more personality. Plus, the volunteers love writing them!

Practical Application

I knew that I couldn’t possibly be the only one struggling with these challenges, so I published a paper sharing these findings with the broader animal welfare community. The paper presents a few case studies to demonstrate how these scales can be applied in real-world shelter settings:

  • Case Study 1: Typical Stress Response and Adaptation

    • A cat exhibiting poor welfare initially, but our data reveal a gradual improvement over two to three days — likely owing to habituation to the shelter environment. These results told us that the base-level enrichment and training plan we were providing was probably sufficient for that particular cat’s needs, once they got used to the initial shock of being somewhere new.
  • Case Study 2: Intervention Success with Gabapentin

    • Another cat exhibiting poor welfare during their first few days at the shelter, but this cat did not show improvement over a similar time frame. On day four, our veterinarian prescribed gabapentin, a psychopharmaceutical. By day five, the cat’s scores dramatically improved. Instead of relying on anecdotal reports, we could quantifiably demonstrate the drug’s effectiveness, which in turn allows veterinarians to confidently prescribe gabapentin for similar cases in the future.
  • Case Study 3: Justifying Alternative Placement

    • This cat came to the shelter with a poor socialization history, but returning him to the field was not possible. While trying to determine the best placement for him, we implemented various intervention techniques, including different psychopharmaceuticals, out-of-cage housing, an appropriate social companion, and formal behavior modification. However, none of these interventions improved the cat’s behavioral indicators of well-being. By relying on the relatively objective data our scales provided, we were able to justify placing this cat in our barn cat program, where he could thrive in a more suitable environment.

The Value of Training

My analysis demonstrated that after initial training, the scales exhibited excellent inter-observer reliability and agreement. But that’s the key — training is required. The Cat-Stress-Score (Kessler & Turner, 1997) is a seven-level ordinal rating scale evaluating 11 different postural/behavioural categories, and it is used ubiquitously in the literature as a method for evaluating how stressed a cat is based on behavioural presentation. I myself have used it in several publications. The trouble is, the main training available for it is a matrix of operational definitions for each level and each postural/behavioural category. It is a fantastic tool, but I have certainly struggled with interpreting exactly how to use it and would have benefitted from a more formal training program. The lack of a training program for the CSS has been lamented in the literature in the past (Finka et al., 2014).

In light of this, I have decided to make my training program available free online to help individuals and organizations learn how to use the scales with confidence. This course provides an overview of the scores, shows multiple video examples of each operationally defined level of each scale, and evaluates the learner’s ability to achieve acceptable levels of inter-observer reliability and agreement, ensuring that they can apply the scales accurately in real-world settings.

Participants in the course are assessed to ensure they meet the reference standard of agreement, and upon successful completion, they are awarded 1 CEU from the Certification Council for Professional Dog Trainers and 1 CEU from the International Association of Animal Behavior Consultants.

The course allows learners to gain a deeper understanding of feline behavior while providing them with a valuable tool for evaluating the welfare of cats in shelters, homes, or other environments where discrete reporting of data is ideal. I hope to one day analyze and share publicly the inter-observer reliability of this wider audience trained in an asynchronous way!

The Feline Behavioural Ordinal Rating Scales Training Course can be accessed from the Toronto Human Society website under ‘courses’.

Applications in Shelters and Beyond

These ordinal rating scales, when paired with the training program, provide an invaluable tool for shelters and other organizations working with cats. These scales can help staff assess and monitor welfare in a consistent and relatively objective way, tracking behavioral progress and identifying cats that may need additional intervention. In shelters, where space and resources are limited, these scales can support decisions about placement, adoption, and potential behavioral modification plans. Additionally, the data collected from these assessments can be used to evaluate the success of novel shelter enrichment programs and determine the effectiveness of behavioral interventions.

Outside of shelters, the scales can be beneficial for veterinary professionals, behavior consultants, and anyone working to improve the well-being of cats in various settings. For behavior consultants, the scales offer a standardized language that can be shared between you and with cat caregivers, facilitating clearer understanding of the presenting behaviours, and allowing more agile adjustments to behavior modification plans with the goal of individualizing care. Veterinary teams can use the scales as a diagnostic tool to track behavioral health over time.

In environments where understanding cat welfare is essential — such as adoption agencies, rescue organizations, and even private homes — these scales can provide a consistent, reliable measure of a cat’s well-being.

Conclusion

The introduction of these ordinal rating scales to assess feline behavior and welfare is a significant advancement in the field of animal behavior. These scales offer a reliable, consistent, and relatively objective method for evaluating the emotional well-being of cats, providing shelter staff and behavior professionals with clear data to guide decisions related to welfare, interventions, and placement.

With excellent inter-observer reliability and agreement demonstrated in the research, these scales are an important tool that can be applied across a wide range of settings, helping to improve the lives of cats in shelters and beyond.

For anyone looking to enhance their skills in using these scales, the free, online training course offers a valuable opportunity to gain certification and practical experience in feline behavior assessment. Whether you are a shelter worker, a behavior consultant, or someone with a deep interest in feline welfare, this training can help you develop the skills needed to use these scales effectively and make a measurable difference in the lives of cats.

References

Ellis, J. J. (2022). Beyond “doing better”: Ordinal rating scales to monitor behavioural indicators of well-being in cats. Animals, 12(21), 2897. https://doi.org/10.3390/ani12212897

Fear Free. N.D. FAS Spectrum and Pain Algorithm. https://fearfreepets.com/fas-spectrum/

Finka, L. R., Ellis, S. L., & Stavisky, J. (2014). A critically appraised topic (CAT) to compare the effects of single and multi-cat housing on physiological and behavioural measures of stress in domestic cats in confined environments. BMC Veterinary Research, 10, 73. https://doi.org/10.1186/1746-6148-10-73

Kessler, M. R., & Turner, D. C. (1997). Stress and Adaptation of Cats (Felis Silvestris Catus) Housed Singly, in Pairs and in Groups in Boarding Catteries. Animal Welfare, 6(3), 243–254. https://doi.org/10.1017/S0962728600019837


Jacklyn Ellis is board certified by the Animal Behavior Society as a Certified Applied Animal Behaviorist (CAAB), is Certified in Shelter Behavior – Cat by the International Association of Animal Behavior Consultants, and is the Director of Behavior at Toronto Humane Society. She earned her PhD in Animal Welfare at the Atlantic Veterinary College, University of Prince Edward Island, where she conducted research on methods for reducing stress in shelter cats. Her work has been published widely in peer reviewed journals and she has presented at many national and international conferences, particularly on feline stress and elimination behavior. She has recently authored two chapters for a new edition of the leading textbook on the behavior and welfare of shelter animals.

TO CITE: Ellis, J. (2025). Beyond “doing better”: Improving the objectivity of cat behavior assessment. IAABC Foundation Journal, 31. doi: https://www.doi.org/10.55736/iaabcfj31.4

SHARE