Services on Demand
Journal
Article
Indicators
- Cited by SciELO
- Access statistics
Related links
- Cited by Google
- Similars in SciELO
- Similars in Google
Share
Revista Facultad Nacional de Salud Pública
Print version ISSN 0120-386XOn-line version ISSN 2256-3334
Abstract
CORREA M, Juan C and VALENCIA C, Marisol. The problem of separation in logistic regression, a solution and an application. Rev. Fac. Nac. Salud Pública [online]. 2011, vol.29, n.3, pp.281-288. ISSN 0120-386X.
Logistic regression is one of the most used statistical techniques for explaining the probabilistic behavior of a given phenomenon. Data separation is a frequent problem in this model, as successes appear separated from failures and make it impossible to find the maximum likelihood estimators. Objective: to present a revision and a solution to the problem, and to compare it with other solutions. METHODOLOGY: a simulation of the logistic model and an estimation of the parameters' bias using the proposed classical and Bayesian solution with fictitious observations, as well as the Firth method. Results: the bias found is lower when the pair of fictitious observations are generated using the Bayesian method. An example about the age at which menarche occurs is presented. DISCUSSION: an appropriate solution to the problem of separation is provided using a simulation in a simple logistic model. CONCLUSIONS: the generation of fictitious observations within the separation region is recommended, and the best solution method is based on Bayesian theory, which achieves convergence of the parameters of the logistic model.
Keywords : logistic model; maximum likelihood estimation; menarche.