Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Hierarchical models for small area estimation using zero-inflated forest inventory variables: comparison and implementation

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

National Forest Inventory (NFI) data are typically limited to sparse networks of sample locations due to cost constraints. While design-based estimators provide reliable forest parameter estimates for large areas, there is increasing interest in model-based small area estimation (SAE) methods to improve precision for smaller spatial, temporal, or biophysical domains. SAE methods can be broadly categorized into area- and unit-level models, with unit-level models offering greater flexibility, making them the focus of this study. Ensuring valid inference requires satisfying model distributional assumptions, which is particularly challenging for NFI variables that exhibit positive support and zero-inflation, such as forest biomass, carbon, and volume. Here, we evaluate nine candidate estimators, including two-stage unit-level hierarchical Bayesian models, single-stage Bayesian models, and two-stage frequentist models, for estimating forest biomass at the county level in Nevada and Washington, United States. Estimator performance is assessed using repeated sampling from simulated populations and unit-level cross-validation with FIA data. Results show that small area estimators incorporating a two-stage approach to account for zero-inflation, county-specific random intercepts and residual variances, and spatial random effects yield the most accurate and well-calibrated county-level estimates, with spatial effects providing the greatest benefits when spatial autocorrelation is present in the underlying population.

Similar Papers
  • Research Article
  • Cite Count Icon 2
  • 10.34123/icdsos.v2021i1.248
Big Data for Small Area Estimation: Happiness Index with Twitter Data
  • Jan 4, 2022
  • Proceedings of The International Conference on Data Science and Official Statistics
  • Sheerin Dahwan Aziz + 1 more

Data availability for small area level is one of the keys to the success of regional development. However, direct estimation of small areas can produce high error due to inadequate sample sizes so the estimation is not reliable. One of alternative solution to this problem is to use the Small Area Estimation (SAE) method which can improve precision by "borrows strength" of the corresponding region information or auxiliary variable information that is strongly related to the response variable. This study uses two SAE models, namely SAE EBLUP Fay-Herriot model with auxiliary variables Podes data and SAE with Error Measurement with auxiliary variable Twitter data. Estimation results using the SAE method are better than direct estimates. This is shown by the RSE value which produced from SAE method, both the EBLUP model and Measurement Error, is smaller than the direct estimate. Therefore, big data can be used as an alternative variable in the SAE model because the data is available in real-time, covers up to the smallest area, and relatively low cost.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 37
  • 10.1371/journal.pone.0189401
Analysis of area level and unit level models for small area estimation in forest inventories assisted with LiDAR auxiliary information
  • Dec 7, 2017
  • PLOS ONE
  • Francisco Mauro + 3 more

Forest inventories require estimates and measures of uncertainty for subpopulations such as management units. These units often times hold a small sample size, so they should be regarded as small areas. When auxiliary information is available, different small area estimation methods have been proposed to obtain reliable estimates for small areas. Unit level empirical best linear unbiased predictors (EBLUP) based on plot or grid unit level models have been studied more thoroughly than area level EBLUPs, where the modelling occurs at the management unit scale. Area level EBLUPs do not require a precise plot positioning and allow the use of variable radius plots, thus reducing fieldwork costs. However, their performance has not been examined thoroughly. We compared unit level and area level EBLUPs, using LiDAR auxiliary information collected for inventorying 98,104 ha coastal coniferous forest. Unit level models were consistently more accurate than area level EBLUPs, and area level EBLUPs were consistently more accurate than field estimates except for large management units that held a large sample. For stand density, volume, basal area, quadratic mean diameter, mean height and Lorey’s height, root mean squared errors (rmses) of estimates obtained using area level EBLUPs were, on average, 1.43, 2.83, 2.09, 1.40, 1.32 and 1.64 times larger than those based on unit level estimates, respectively. Similarly, direct field estimates had rmses that were, on average, 1.37, 1.45, 1.17, 1.17, 1.26, and 1.38 times larger than rmses of area level EBLUPs. Therefore, area level models can lead to substantial gains in accuracy compared to direct estimates, and unit level models lead to very important gains in accuracy compared to area level models, potentially justifying the additional costs of obtaining accurate field plot coordinates.

  • Research Article
  • Cite Count Icon 1
  • 10.1016/j.foreco.2025.122999
Leveraging national forest inventory data to estimate forest carbon density status and trends for small areas
  • Nov 1, 2025
  • Forest Ecology and Management
  • Elliot S Shannon + 7 more

Leveraging national forest inventory data to estimate forest carbon density status and trends for small areas

  • Research Article
  • Cite Count Icon 4
  • 10.3141/2105-09
Model-Based Synthesis of Household Travel Survey Data in Small and Midsize Metropolitan Areas
  • Jan 1, 2009
  • Transportation Research Record: Journal of the Transportation Research Board
  • Liang Long + 2 more

Household travel data synthesis–simulation has become a promising alternative or supplement to survey data from both small urban areas and large metropolitan regions in which data are expensive to collect or the data required to support the planning process have become outdated. This paper proposes and applies model-based approaches [i.e., small area estimation (SAE) methods] to synthesize household travel characteristics. The proposed methods address the sampling-bias concerns in the existing methods. Specifically, three SAE methods–-the generalized regression estimators method, the empirical best linear unbiased predictor (EBLUP) method, and the synthetic method (an EBLUP without random area effects)–-are applied to synthesize household travel characteristics at both census tract and individual levels. The SAE framework of synthesizing household travel characteristics is demonstrated with the National Household Travel Survey data and the Census Transportation Planning Package data in the Des Moines metropolitan area in central Iowa. Results indicate that SAE methods are promising approaches to synthesize unbiased aggregate and disaggregate household travel characteristics by incorporating population auxiliary information and local, small-household travel survey data. The proposed data synthesis methods and analysis findings will provide a useful tool for practitioners, planners, and policy makers in transportation analyses. The paper also points out that by linking population synthesis with the travel data simulation framework described here, this method could be of broad application in transportation planning.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 11
  • 10.3389/ffgc.2022.745874
Simplifying Small Area Estimation With rFIA: A Demonstration of Tools and Techniques
  • Apr 12, 2022
  • Frontiers in Forests and Global Change
  • Hunter Stanke + 2 more

The United States (US) Department of Agriculture Forest Service Forest Inventory and Analysis (FIA) program operates the national forest inventory of the US. Traditionally, the FIA program has relied on sample-based approaches—permanent plot networks and associated design-based estimators—to estimate forest variables across large geographic areas and long periods of time. These approaches generally offer unbiased inference on large domains but fail to provide reliable estimates for small domains due to low sample sizes. Rising demand for small domain estimates will thus require the FIA program to adopt non-traditional estimation approaches that are capable of delivering defensible estimates of forest variables at increased spatial and temporal resolution, without the expense of collecting additional field data. In light of this challenge, the development of small area estimation (SAE) methods—estimation techniques that support inference on small domains—for FIA data has become an active and highly productive area of research. Yet, SAE methods remain difficult to apply to FIA data, due in part to the complex data structures and survey design used by the FIA program. Herein, we present the potential of rFIA, an open-source R package designed to increase the accessibility of FIA data, to simplify the application of a broad suite of SAE methods to FIA data. We demonstrate this potential via two case studies: (1) estimation of contemporary county-level forest carbon stocks across the conterminous US using a spatial Fay-Herriot model; and (2) temporally-explicit estimation of multi-decadal trends in merchantable wood volume in Washington County, Maine using a Bayesian multi-level model. In both cases, we show the application of SAE techniques offers considerable improvements in precision over FIA's traditional, post-stratified estimators. Finally, we offer a discussion of the potential role that rFIA and other open-source tools might play in accelerating the adoption of SAE techniques among users of FIA data.

  • Research Article
  • 10.71014/sieds.v79i3.345
Small Area Estimation of Poverty Indicators
  • Feb 28, 2025
  • Rivista Italiana di Economia Demografia e Statistica
  • Michele D'Alò + 3 more

ISTAT has been carrying out extensive research to implement Small Area Estimation (SAE) methods for computing Sustainable Development Goals (SDGs) indicators related to health, occupational status, gender equality, and poverty. This work aims to present the main results obtained applying some SAE methods to estimate the "At Risk of Poverty" indicator for unplanned domains using EU-SILC data. The sub-domains of interest are the provinces (NUTS3) and metropolitan cities, while the survey is designed to provide estimates up to the NUTS2 level (regions). The Small Area Estimation (SAE) methods considered encompass both area and unit-level mixed models, and their results are compared against each other. Administrative data sourced from ISTAT's Integrated System of Registers (ISR), specifically from the Population Register and the Labour Register, integrated with income-related administrative data, are used to specify the models. Furthermore, with direct estimates and administrative auxiliary information available from 2017 to 2021, SAE methods can borrow strength not only from other areas but also from various survey cycles. A final step in the process of estimating small-area statistics through an inferential model-based approach is establishing coherence between estimations of the target indicator computed at various levels of granularity. It is performed to align SAEs with precise and unbiased direct estimates computed at higher planned domain levels. This final calibration is not merely cosmetic. It is essential to meet user requirements on coherence and also to enhance the overall accuracy and reliability of model-based SAEs. The application of Small Area Estimation (SAE) estimates allows gains of efficiency compared to direct estimates.

  • Dissertation
  • 10.20378/irb-103063
Recent Advances in Small Area Estimation of Economic and Poverty Indicators using Traditional and Alternative Data
  • Jan 1, 2024
  • Yeonjoo Lee

Chapter 1 - Variable selection using conditional AIC for linear mixed models with data-driven transformations When data analysts use linear mixed models, they usually encounter two practical problems: a) the true model is unknown and b) the Gaussian assumptions of the errors do not hold. While these problems commonly appear together, researchers tend to treat them individually by a) finding an optimal model based on the conditional Akaike information criterion (cAIC) and b) applying transformations on the dependent variable. However, the optimal model depends on the transformation and vice versa. In this paper, we aim to solve both problems simultaneously. In particular, we propose an adjusted cAIC by using the Jacobian of the particular transformation such that various model candidates with differently transformed data can be compared. From a computational perspective, we propose a step-wise selection approach based on the introduced adjusted cAIC. Model-based simulations are used to compare the proposed selection approach to alternative approaches. Finally, the introduced approach is applied to Mexican data to estimate poverty and inequality indicators for 81 municipalities. Chapter 2 - Estimation of the consumer price index with regional weights using small area estimation methods: A case study of Germany The consumer price index (CPI) is an important indicator for formulating effective policies related to wage and inflation control. Most countries regularly produce the national CPI and some countries also publish the CPI at sub-national levels. This sub-national (or regional) CPI depicts region specific information and consequently is a helpful tool for local policy makers. Germany provides this regional CPI for all 16 states with national product weights. Still, national products weights do not adequately represent the importance of products at the regional level, whereas a better regional CPI uses regional product weights to reflect regional specifics more accurately. In this study, I explore the estimation of such a better regional CPI with regional weights by using an accessible income and consumption survey from Germany. To obtain reliable regional weights, I focus on estimating regional expenditures for each product. Regional weights of each product are derived by calculating the proportion of the estimated specific product expenditure over the total expenditure. However, estimating regional expenditures is challenging because of the small sample size in each region. Small sample sizes potentially produce unreliable estimates. To address this problem, I propose the use of a small area estimation approach based on multivariate Fay-Herriot (MFH) models. By using MFH models, I show how model-based estimation of regional expenditure improves the reliability of the estimation using the case of Germany and further discuss the limitations as well as future research directions. Chapter 3 - Small area estimation using geospatial data based on transformed two-fold nested error regression models Fighting poverty starts with identifying where exactly poverty is the most severe by estimating poverty indicators. For a precise estimation of indicators at disaggregated regional levels, a small area estimation (SAE) approach is essential. SAE methods require a survey and a auxiliary dataset. By combining two datasets, SAE methods overcome the problem of small sample sizes and produce reliable estimates at a small area level. A census dataset is commonly used as auxiliary data to estimate poverty indicators. However, data protection laws often impede access and even a recent census may already be outdated because in developing countries, changes are as quick as they are dynamic. Therefore, an older census is inappropriate as auxiliary data. Geospatial data is a viable alternative since it is freely available, up to date, and covers all inhabited areas. Still, a central challenge remains: geospatial data is collected at grid level which is usually larger than a household but smaller than a small area. Since traditional SAE models use either a household level model or an area level model, we need to investigate which SAE model optimizes the advantages of grid level data. We suggest the two-fold nested regression model as it allows two random effects at different regional levels. More random effects capture the hierarchical data structure more accurately than standard SAE models. Additionally, we introduce transformations to the two-fold model to correct the violation in distributional model assumptions and to estimate ratio type indicators. We furthermore propose an estimation method for mean squared errors (MSE) with the transformed two-fold model. We show an efficiency gain of the two-fold model compared to the standard Fay-Herriot model, and the proposed MSE estimator is reasonable. Lastly, we apply the proposed transformed two-fold model to the geospatial data of Mozambique to estimate poverty indicators for 161 districts.

  • Book Chapter
  • 10.1007/978-981-15-1476-0_15
Small Area Estimation for Skewed Semicontinuous Spatially Structured Responses
  • Jan 1, 2020
  • Chiara Bocci + 3 more

When surveys are not originally designed to produce estimates for small geographical areas, some of these domains can be poorly represented in the sample. In such cases, model-based small area estimators can be used to improve the accuracy of the estimates by borrowing information from other sub-populations. Frequently, in surveys related to agriculture, forestry or the environment, we are interested in analyzing continuous variables which are characterized by a strong spatial structure, a skewed distribution and a point mass at zero. In such cases, standard methods for small area estimation, which are based on linear mixed models, can be inefficient. The aim of this chapter is to discuss small area estimation models suggested in literature to handle zero-inflated, skewed, spatially structured data and to present them under the unified approach of generalized two-part random effects models.

  • Research Article
  • Cite Count Icon 39
  • 10.1146/annurev-statistics-031219-041212
Robust Small Area Estimation: An Overview
  • Mar 9, 2020
  • Annual Review of Statistics and Its Application
  • Jiming Jiang + 1 more

A small area typically refers to a subpopulation or domain of interest for which a reliable direct estimate, based only on the domain-specific sample, cannot be produced due to small sample size in the domain. While traditional small area methods and models are widely used nowadays, there have also been much work and interest in robust statistical inference for small area estimation (SAE). We survey this work and provide a comprehensive review here. We begin with a brief review of the traditional SAE methods. We then discuss SAE methods that are developed under weaker assumptions and SAE methods that are robust in certain ways, such as in terms of outliers or model failure. Our discussion also includes topics such as nonparametric SAE methods, Bayesian approaches, model selection and diagnostics, and missing data. A brief review of software packages available for implementing robust SAE methods is also given.

  • Research Article
  • Cite Count Icon 101
  • 10.1177/0022034516629112
Predicting Periodontitis at State and Local Levels in the United States
  • Feb 4, 2016
  • Journal of Dental Research
  • P.I Eke + 7 more

The objective of the study was to estimate the prevalence of periodontitis at state and local levels across the United States by using a novel, small area estimation (SAE) method. Extended multilevel regression and poststratification analyses were used to estimate the prevalence of periodontitis among adults aged 30 to 79 y at state, county, congressional district, and census tract levels by using periodontal data from the National Health and Nutrition Examination Survey (NHANES) 2009–2012, population counts from the 2010 US census, and smoking status estimates from the Behavioral Risk Factor Surveillance System in 2012. The SAE method used age, race, gender, smoking, and poverty variables to estimate the prevalence of periodontitis as defined by the Centers for Disease Control and Prevention/American Academy of Periodontology case definitions at the census block levels and aggregated to larger administrative and geographic areas of interest. Model-based SAEs were validated against national estimates directly from NHANES 2009–2012. Estimated prevalence of periodontitis ranged from 37.7% in Utah to 52.8% in New Mexico among the states (mean, 45.1%; median, 44.9%) and from 33.7% to 68% among counties (mean, 46.6%; median, 45.9%). Severe periodontitis ranged from 7.27% in New Hampshire to 10.26% in Louisiana among the states (mean, 8.9%; median, 8.8%) and from 5.2% to 17.9% among counties (mean, 9.2%; median, 8.8%). Overall, the predicted prevalence of periodontitis was highest for southeastern and southwestern states and for geographic areas in the Southeast along the Mississippi Delta, as well as along the US and Mexico border. Aggregated model-based SAEs were consistent with national prevalence estimates from NHANES 2009–2012. This study is the first-ever estimation of periodontitis prevalence at state and local levels in the United States, and this modeling approach complements public health surveillance efforts to identify areas with a high burden of periodontitis.

  • Book Chapter
  • Cite Count Icon 1
  • 10.1007/978-3-319-05320-2_2
M-Quantile Small Area Models for Measuring Poverty at a Local Level
  • Jan 1, 2014
  • Monica Pratesi

M-quantile small area estimation (SAE) methods constitute a set of advanced statistical inference techniques that can be used for the measurement of poverty and living conditions by survey practitioners, researchers in private and public organizations, official statistical agencies, and local governmental agencies. In particular, the estimates produced using these SAE methods are well suited to mapping geographical variations in these conditions. In this paper, we summarize the ideas set out in some recent papers on M-quantile methods and their extensions and also comment on important issues that arise when SAE methods are used in poverty assessment in three Italian Regions.KeywordsM-quantile modelsPoverty mappingSmall area estimation

  • Research Article
  • 10.12962/j20882033.v23i1.13
Empirical Bayesian Method for the Estimation of Literacy Rate at Sub-district Level Case Study: Sumenep District of East Java Province
  • Feb 1, 2012
  • IPTEK The Journal for Technology and Science
  • A.Tuti Rumiati + 3 more

This paper discusses Bayesian Method of Small Area Estimation (SAE) based on Binomial response variable. SAE method being developed to estimate parameter in small area due to insufficiency of sample. The case study is literacy rate estimation at sub-district level in Sumenep district, East Java Province. Literacy rate is measured by proportion of people who are able to read and write, from the population of 10 year-old or more. In the case study we used Social Economic Survey (Susenas)data collected by BPS. The SAE approach was applied since the Susenas data is not representative enough to estimate the parameters at sub-district level because it’s designed to estimate parameters in regional area (in scope of a district/city at minimum). In this research, the response variable being used was logit function trasformation of pi (the parameter of Binomial distribution). We applied direct and indirect approach for parameter estimation, both using Empirical Bayes approach. For direct estimation we used prior distribution of Beta distribution and Normal prior distribution for logit function (pi) and to estimate parameter by using numerical method, i.e integration Monte Carlo. For indirect approach, we used auxiliary variables which are combinations of sex and age (which is divided into five categories). Penalized Quasi Likelihood (PQL) was used to get parameter estimation of SAE model and Restricted Maximum Likelihood method (REML) for MSE estimation. Instead of Bayesian approach, we are also conducting direct estimation using classical approach in order to evaluate the quality of the estimators. This research gives some findings, those are: Bayesian approach for SAE model gives the best estimation because having the lowest MSE value compares to the other methods. For the direct estimation, Bayesian approach using Beta and logit Normal prior distribution give a very similar result to the direct estimation with classical approach since the weight of is too large, which is about 0.905. It is also found that direct estimation using Bayesian approach with the Beta prior distribution gives better MSE than using logit normal prior distribution.

  • Research Article
  • 10.30598/barekengvol18iss2pp1009-1022
TWOFOLD SUBAREA MODEL FOR ESTIMATING COMMUTER PROPORTION IN 10 METROPOLITAN AREAS
  • May 25, 2024
  • BAREKENG: Jurnal Ilmu Matematika dan Terapan
  • Yudi Fathul Amin + 2 more

The metropolitan area is a major contributor to national GDP. The metropolitan area is a center of attraction for many people who come to earn income as commuters. Commuters are people who carry out work activities in the center of the metropolitan area, which are carried out by residents who live in suburban areas around the center of the metropolitan area and commute regularly every day. The availability of commuter statistics from surveys for presentation level down to the smallest administrative level, such as regencies/municipalities, is unreliable. This happens because this level of presentation has poor precision due to insufficient samples due to the Statistics Indonesia survey design for making estimates at the national and provincial levels. It can be done using small area estimation (SAE) to meet increasing data needs, but existing SAE models can often estimate only at one level. To meet data requests more effectively, a model is needed that can estimate several small areas simultaneously. In SAE, one of the SAE models that can do this is the twofold subarea model. The twofold subarea model produces estimates of the proportion of commuters with good precision at the subarea level (regencies/municipalities) and area level (metropolitan area), with the RRMSE percentage value of the estimated proportion of commuters being below 25% for all regions. The results of this research can be used to present commuter data at the regencies/municipalities level and metropolitan area level where there is a lack of samples and become a new opportunity for Statistics Indonesia to increase statistical production in small areas, which is more effective compared to other SAE methods which have so far been used only to estimate one area level.

  • Research Article
  • 10.32628/ijsrset21841116
Comparison of EBLUP and EBLUP Modification in Estimating Small Areas (Study : Percentages of Poverty in Bogor District)
  • Dec 4, 2018
  • International Journal of Scientific Research in Science, Engineering and Technology
  • Hary Merdeka + 2 more

A small area of the sample occurs when the sample size is very small. A large error will get if the parameters estimation is done with small the sample. One method to overcome it using a small area estimation (SAE) method. A small area estimator is a statistical technique to estimate the parameters of a sub-population with a small sample size. Estimates in the small area estimator method is based on the model and are indirect estimates. In this study the indirect method used is the EBLUP method and the modification of EBLUP estimator. The results of the alleged percentage of poverty in the Bogor district show that the EBLUP modification method is better compared to the expected method directly. This is based on the average of the RRMSE obtained.

  • Research Article
  • Cite Count Icon 22
  • 10.1093/forestry/cpz073
A novel application of small area estimation in loblolly pine forest inventory
  • Dec 24, 2019
  • Forestry: An International Journal of Forest Research
  • P Corey Green + 3 more

Loblolly pine (Pinus taeda L.) is one of the most widely planted tree species globally. As the reliability of estimating forest characteristics such as volume, biomass and carbon becomes more important, the necessary resources available for assessment are often insufficient to meet desired confidence levels. Small area estimation (SAE) methods were investigated for their potential to improve the precision of volume estimates in loblolly pine plantations aged 9–43. Area-level SAE models that included lidar height percentiles and stand thinning status as auxiliary information were developed to test whether precision gains could be achieved. Models that utilized both forms of auxiliary data provided larger gains in precision compared to using lidar alone. Unit-level SAE models were found to offer additional gains compared with area-level models in some cases; however, area-level models that incorporated both lidar and thinning status performed nearly as well or better. Despite their potential gains in precision, unit-level models are more difficult to apply in practice due to the need for highly accurate, spatially defined sample units and the inability to incorporate certain area-level covariates. The results of this study are of interest to those looking to reduce the uncertainty of stand parameter estimates. With improved estimate precision, managers, stakeholders and policy makers can have more confidence in resource assessments for informed decisions.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant