Skip to content
Research Article Open access CC BY 3.0

Models for Injury Count Data in the U.S. National Health Interview Survey

Jin Peng, Tianmeng Lyu, Junxin Shi, Haikady N. Nagaraja, Huiyun Xiang

Journal of Scientific Research and Reports · pp. 2286–2302 · Published 17 Jul 2014

10.9734/JSRR/2014/9490

Abstract

Aims: To examine the best count data model for injury data in the National Health Interview Survey (NHIS). To compare the best count data model with traditional logistic regression model in analyzing injury data in NHIS. Data Source: 2006-2010 medically consulted non-occupational injury data from National Health Interview Survey (NHIS). Methodology: Six count data models (Poisson, negative binomial (NB), zero-inflated Poisson (ZIP), zero-inflated NB (ZINB), hurdle Poisson (HP), and hurdle NB (HNB)) were compared using Likelihood Ratio (LR) test and Vuong test. Injury count was used as the dependent variable in count data models. Independent variables included age, gender, marital status, race, education, poverty status, disability status and medical insurance coverage status. Dichotomized injury count was used as the dependent variable in logistic regression model. The same independent variables used in count data models were included in logistic regression model. The model fit of logistic regression was examined by Hosmer and Lemeshow goodness of fit test. Results: Among 248,850 participants aged 18-64, 98.37% have no medically consulted non-occupational injuries, 1.55% have 1 medically consulted non-occupational injury, 0.07% have 2 or more medically consulted non-occupational injuries. Zero-inflated negative binomial (ZINB) model offered the best fit. Logistic regression model provided a good fit but resulted in different estimates from ZINB model. Conclusion: Zero-inflated negative binomial (ZINB) model demonstrated the potential to be the best model for injury count data with excess zeros. Given the infrequent occurrence of multiple injuries in our data, the logistic regression model is appropriate for assessing injury burden and identifying injury risks. However, for more frequently-occurring injuries (e.g. sports injuries), logistic regression may undercount the total number of injuries and result in biased estimates. The evaluation procedure and model selection criteria presented in this paper provide a useful approach to modeling injury count data with excess zeros.

Injury epidemiology National Health Interview Survey (NHIS) injury count data logistic regression model Zero-inflated Negative Binomial (ZINB) model

Cited by 2

SAYMA VERİLERİNİN MODELLENMESİ VE BİREYLERİN İŞSİZ KALMA SÜRESİ ÜZERİNE BİR UYGULAMA

Afet SÖZEN ÖZDEN, Elvan HAYAT · Pamukkale University Journal of Social Sciences Institute · 2022

Key Factors Influencing the Incidence of West Nile Virus in Burleigh County, North Dakota

Hiroko Mori, Joshua Wu, Motomu Ibaraki · International Journal of Environmental Research and Public Health · 2018

Article metrics

Real usage data collected on this platform.

0

Page views

0

PDF downloads

0

Outbound clicks

2

Citations

Views by country

Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".

No views recorded yet.

Traffic sources

Referring site, by host.

No traffic recorded yet.

Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.