Hostname: page-component-76d6cb85b7-ntvhh Total loading time: 0 Render date: 2026-07-24T00:02:54.766Z Has data issue: false hasContentIssue false

A hurdle model for ordinal scoring data with an underlying percentage scale

Published online by Cambridge University Press:  12 February 2026

Emilia Koch*
Affiliation:
Biostatistics Unit, Institute of Crop Science, University of Hohenheim, Germany
Jens Hartung
Affiliation:
Department Sustainable Agriculture and Energy Systems, University of Applied Science, Weihenstephan-Triesdorf, Germany
Til Feike
Affiliation:
Institute for Strategies and Technology Assessment, Julius Kühn Institute Kleinmachnow, Germany
Benjamin Epler
Affiliation:
AGROTO GmbH, Germany
Hans-Peter Piepho
Affiliation:
Biostatistics Unit, Institute of Crop Science, University of Hohenheim, Germany
*
Corresponding author: Emilia Koch; Email: emiliakoch@gmx.net
Rights & Permissions [Opens in a new window]

Abstract

It is crucial to properly evaluate the traits that directly impact agricultural productivity. Some of these traits, such as soil erosion or crop diseases, are quantified with scoring systems. The resulting data are strictly ordinal and often have an underlying percentage scale. Deciding which model to use for this type of data is not straightforward. Ordinal scores do not meet the assumptions required for analysis of variance. Although multinomial ordinal models, particularly the threshold model, can be applied, they do not account for the underlying percentage scale of the data. To address this limitation, a hurdle model tailored for interval-censored percentage data is proposed. It is a two-part model that models the data according to their nature: In its first part, it models presence or absence of a disease (incidence), and in the second part it models severity or abundance. Individually modelling presence and absence in the first part allows to account for zero inflation. The second part implements theory from the threshold model and the Johnson SB system of distributions that involves a transformation of the percentage scale to a normal distribution. The model result also reflects the two components. They individually describe the degree of disease infestation, and the degree of disease spread. This improves interpretability and enables concrete, insightful conclusions. To illustrate the model, mildew scorings from an on-farm trial in grapevines were used. The model was found highly suitable for this dataset and superior to the threshold model.

Information

Type
Crops and Soils Research Paper
Creative Commons
Creative Common License - CCCreative Common License - BY
This is an Open Access article, distributed under the terms of the Creative Commons Attribution licence (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted re-use, distribution and reproduction, provided the original article is properly cited.
Copyright
© The Author(s), 2026. Published by Cambridge University Press
Figure 0

Figure 1. Example of an ordinal rating scale with an underlying percentage scale for disease ratings. Each ordinal score represents a defined range of percentages. Thresholds are the limits that indicate the transition to a new score. For this scale, there are five thresholds and the point score of zero.

Figure 1

Table 1. Overview of the steps for analysing the data with the hurdle model, including the binary part (INCIDENCE) and the continuous part (SEVERITY). The linear predictor is defined on the logit scale, which ranges from −∞ to ∞. On this scale, terms combine additively. The probability is then obtained by applying the inverse logit

Figure 2

Figure 2. Illustration of the hurdle model, starting from the left with disease scorings, continuing with the binary model component (INCIDENCE), and the continuous component (SEVERITY). A scoring scheme with seven classes $\left( {k = 1,{\rm{\;}} \ldots, {\rm{\;}}7} \right)$ was used. INCIDENCE’s probability of $k = 1$, of being disease free, is ${p_{\rm{1}}}$. The probabilities for disease severity for $k = 2,{\rm{\;}} \ldots, {\rm{\;}}7$ are ${q_2}$ to ${q_{\rm{7}}}.$ These are provided by the area under the normal density function and are separated by the thresholds ${\theta _2}$ to ${\theta _6}$ (Piepho and Kalka, 2003).

Figure 3

Table 2. Overview table for parameters and indexes

Figure 4

Table 3. Comparing the hurdle model with separate linear predictors with the hurdle model with linearly connected linear predictors and the threshold model. Models were compared based on the Akaike Information Criterion (AIC). A smaller AIC indicates a better model fit. Each combination of the observed plant organ and year was evaluated as own dataset. Each dataset was evaluated with all three models. For better readability, the smallest AIC within each dataset is highlighted in bold

Figure 5

Table 4. Results of significance tests for the hurdle model with independent linear predictors

Figure 6

Figure 3. Mean mildew infestation [%] on leaves for data sets with a significant treatment effect. Treatments (T_1 to T_8) with the same letter are not significantly different. The error bar shows the 95% confidence limits. INCIDENCE = result of the first part of the hurdle model; Disease incidence = probability of being infected × 100%; SEVERITY = result of the second part of the hurdle model; Disease severity = area of infected plant tissue (in %).

Figure 7

Figure 4. Mean mildew infestation [%] on fruits for data sets showing a significant treatment effect. The error bar shows the 95% confidence limits. Treatments with at least one identical letter are not significantly different. INCIDENCE = result of the first part of the hurdle model; Disease incidence = probability of being infected × 100%; SEVERITY = result of the second part of the hurdle model; Disease severity = area of infected plant tissue (in %).

Supplementary material: File

Koch et al. supplementary material 1

Koch et al. supplementary material
Download Koch et al. supplementary material 1(File)
File 198.6 KB
Supplementary material: File

Koch et al. supplementary material 2

Koch et al. supplementary material
Download Koch et al. supplementary material 2(File)
File 13.5 KB
Supplementary material: File

Koch et al. supplementary material 3

Koch et al. supplementary material
Download Koch et al. supplementary material 3(File)
File 938.6 KB
Supplementary material: File

Koch et al. supplementary material 4

Koch et al. supplementary material
Download Koch et al. supplementary material 4(File)
File 176.8 KB