Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Diagnostic test for the circular-circular additive regression model

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Diagnostic test for the circular-circular additive regression model

Similar Papers
  • Research Article
  • Cite Count Icon 1
  • 10.1088/1742-6596/1524/1/012124
Probability of caused factors stroke disease use link and reliability functions
  • Apr 1, 2020
  • Journal of Physics: Conference Series
  • Sudarno + 2 more

Diabetes mellitus disease is disease which abnormal metabolism for a long time, because pancreas can not be able to produce insulin hormone be enough, or because body can not be able to use insulin hormone has been produced by effective. A stroke occurs if the flow of oxygen-rich blood to a portion of the brain is blocked. Without oxygen, brain cells start to die after a few minutes. Sudden bleeding in the brain also can cause a stroke if it damages brain cells. These objectives are finding significant factors which cause diabetes mellitus disease and determine ordinal regression model. Ordinal regression model is used to look for probability and reliability functions of a patient has stroke disease. The method used to three link functions, that are logit link function, normit link function, and cloglog link function. Testing of homogeneity prediction result of link functions uses linear hypothesis test. Factors caused diabetes mellitus are body mass index, high density lipoprotein, and albuminuria. These factors cause to diabetes mellitus and stroke could be used to prevent diseases, in order to all persons are healthy and happy. The result that probability of a patient with macroalbuminuria has stroke greater than microalbuminuria and a patient with microalbuminuria has stroke greater than normal. Probability of patient with macroalbuminuria by logit, normit, and clogloc link functions is decrease, respectively. Probability of patient with microalbuminuria by logit, normit, and cloglog link functions is increase, respectively. Reliability of a patient with macroalbuminuria, normal, and microalbuminuria have stroke, respectively, is decrease. Reliability of patient with macroalbuminuria by logit, normit, and clogloc link functions, respectively, is increase. Reliability of patient with microalbuminuria by logit, normit, and clogloc link functions, respectively, is zero. All of link function methods yield estimation probability value is the same. AIC value of logit link function, normit link function, and cloglog link function are, respectively, 167.6826, 168.3965, and 169.6107. These results are same by the result of linear hypothesis analysis that AIC values are not different meanwhile their AIC values are not equal. Therefore, logit model, normit model and cloglog model could be used to predict probability with result almost same.

  • Research Article
  • Cite Count Icon 3
  • 10.1007/s11222-017-9781-3
A new algorithm to estimate monotone nonparametric link functions and a comparison with parametric approach
  • Oct 20, 2017
  • Statistics and Computing
  • Xin Wang + 2 more

The generalized linear model (GLM) is a class of regression models where the means of the response variables and the linear predictors are joined through a link function. Standard GLM assumes the link function is fixed, and one can form more flexible GLM by either estimating the flexible link function from a parametric family of link functions or estimating it nonparametically. In this paper, we propose a new algorithm that uses P-spline for nonparametrically estimating the link function which is guaranteed to be monotone. It is equivalent to fit the generalized single index model with monotonicity constraint. We also conduct extensive simulation studies to compare our nonparametric approach for estimating link function with various parametric approaches, including traditional logit, probit and robit link functions, and two recently developed link functions, the generalized extreme value link and the symmetric power logit link. The simulation study shows that the link function estimated nonparametrically by our proposed algorithm performs well under a wide range of different true link functions and outperforms parametric approaches when they are misspecified. A real data example is used to illustrate the results.

  • Research Article
  • Cite Count Icon 598
  • 10.1214/009053604000001156
Generalized functional linear models
  • Apr 1, 2005
  • The Annals of Statistics
  • Hans-Georg Müller + 1 more

We propose a generalized functional linear regression model for a regression situation where the response variable is a scalar and the predictor is a random function. A linear predictor is obtained by forming the scalar product of the predictor function with a smooth parameter function, and the expected value of the response is related to this linear predictor via a link function. If, in addition, a variance function is specified, this leads to a functional estimating equation which corresponds to maximizing a functional quasi-likelihood. This general approach includes the special cases of the functional linear model, as well as functional Poisson regression and functional binomial regression. The latter leads to procedures for classification and discrimination of stochastic processes and functional data. We also consider the situation where the link and variance functions are unknown and are estimated nonparametrically from the data, using a semiparametric quasi-likelihood procedure. An essential step in our proposal is dimension reduction by approximating the predictor processes with a truncated Karhunen–Loève expansion. We develop asymptotic inference for the proposed class of generalized regression models. In the proposed asymptotic approach, the truncation parameter increases with sample size, and a martingale central limit theorem is applied to establish the resulting increasing dimension asymptotics. We establish asymptotic normality for a properly scaled distance between estimated and true functions that corresponds to a suitable L2 metric and is defined through a generalized covariance operator. As a consequence, we obtain asymptotic tests and simultaneous confidence bands for the parameter function that determines the model. The proposed estimation, inference and classification procedures and variants with unknown link and variance functions are investigated in a simulation study. We find that the practical selection of the number of components works well with the AIC criterion, and this finding is supported by theoretical considerations. We include an application to the classification of medflies regarding their remaining longevity status, based on the observed initial egg-laying curve for each of 534 female medflies.

  • Research Article
  • Cite Count Icon 2
  • 10.5430/wje.v3n4p96
Identifying the Factors that Influence Change in SEBD Using Logistic Regression Analysis
  • Aug 19, 2013
  • World Journal of Education
  • Liberato Camilleri + 1 more

Multiple linear regression and ANOVA models are widely used in applications since they provide effective statistical tools for assessing the relationship between a continuous dependent variable and several predictors. However these models rely heavily on linearity and normality assumptions and they do not accommodate categorical dependent variables. The seminal contribution of John Nelder and Robert Wedderburn (1972) introduced the concept of Generalized Linear Models. GLMs overcome the limitations of Normal regression models and accommodate any distribution which is a member of the exponential family. Moreover, these models relate the dependent variable to the linear predictor (non-random component) through any invertible link function. Logistic regression models are GLMs that accommodate categorical dependent variables. They assume a Binomial distribution and Logit canonical link function. The iteratively re-weighted least squares algorithm using the Fisher scoring technique is employed to maximize the log-likelihood function in GLMs and estimate the model parameters. In this paper, Logistic regression analysis was used to identify the dominant factors that influence change in social, emotional and behaviour difficulties (SEBD) of Maltese children. The study comprised 486 pupils whose SEBD was assessed by both teachers and parents using the Strengths and Difficulties Questionnaire (Goodman 1997) when the children were aged 6 and 9 years old.

  • Supplementary Content
  • 10.17877/de290r-390
Regression and residual analysis in linear models with interval censored data
  • Dec 14, 2004
  • Technische Universität Dortmund Eldorado (Technische Universität Dortmund)
  • Rebekka Topp

This work consists of two parts, both related with regression analysis for interval censored data. Interval censored data x have the property that their value cannot be observed exactly but only the respective interval [xL,xR] which contains the true value x with probability one. In the first part of this work I develop an estimation theory for the regression parameters of the linear model where both dependent and independent variables are interval censored. In doing so I use a semi-parametric maximum likelihood approach which determines the parameter estimates via maximization of the likelihood function of the data. Since the density function of the covariate is unknown due to interval censoring, the maximization problem is solved through an algorithm which frstly determines the unknown density function of the covariate and then maximizes the complete data likelihood function. The unknown covariate density is hereby determined nonparametrically through a modification of the approach of Turnbull (1976). The resulting parameter estimates are given under the assumption that the distribution of the model errors belong to the exponential familiy or are Weibull. In addition I extend my extimation theory to the case that the regression model includes both an interval censored and an uncensored covariate. Since the derivation of the theoretical statistical properties of the developed parameter estimates is rather complex, simulations were carried out to determine the quality of the estimates. As a result it can be seen that the estimated values for the regression parameters are always very close the real ones. Finally, some alternative estimation methods for this regression problem are discussed. In the second part of this work I develop a residual theory for the linear regression model where the covariate is interval censored, but the depending variable can be observed exactly. In this case the model errors appear to be interval censored, and so the residuals. This leads to the problem of not directly observable residuals which is solved in the following way: Since one assumption of the linear regression model is the N(0,2)-distribution of the model errors, it follows that the distribtuion of the interval censored errors is a truncated normal distribution, the truncation being determined by the observed model error intervals. Consequently, the distribution of the interval censored residuals is a -distribution, truncated in the respective residual interval, where the estimation of the residual variance is accomplished through the method of Gomez et al. (2002). In a simulation study I compare the behaviour of the so constructed residuals with those of Gomez et al. (2002) and a naive type of resiudals which considers the middle of the residual interval as the observed residual. The results show that my residuals can be used for most of the simulated scenarios, wheras this is not the case for the other two types of residuals. Finally, my new residual theory is applied to a data set from a clinical study.

  • Book Chapter
  • Cite Count Icon 3
  • 10.1002/9781118445112.stat06576
Generalized Linear Models: Introduction
  • Sep 29, 2014
  • Wiley StatsRef: Statistics Reference Online
  • Brian S Everitt

Generalized linear models provide a unified framework for regression models including multiple regression, logistic regression, analysis of variance, and analysis of covariance. Such models consist of three components – a link function specifying the transformation of the response variable to be modeled by a linear function of the explanatory variables, an error distribution suitable for each type of response, and a variance function specifying the relationship between the mean and variance of the error distribution.

  • Research Article
  • 10.11648/j.bsi.20170204.13
On the Comparison of Some Link Functions of Binary Response Analysis Under Symmetric and Asymmetric Assumptions
  • Sep 23, 2017
  • Biomedical Statistics and Informatics
  • Saddam Adams Damisa + 5 more

Binary response analysis is modeled when the response variable is nominal and as such violates the use of the ordinary linear regression model. This paper utilizes the classical approach to fit a categorical response regression model using the logit, probit. loglog and the complementary loglog (Cloglog) link functions under symmetric and asymmetric assumptions. It is captured in past studies that we can only make comparisons between these link functions when n is large say (n > 1000), In this study we compared the link functions to investigate this claim with small values of n less than 1000. We fit the Cloglog and loglog models on 600 tuberculosis patients who may be co-infected with hypertension while the R package was initiated in simulating a binary data for fitting the logit and probit models using the Akaike Information Criterion (AIC) as a basis of comparison for the symmetric and asymmetric different model fitting techniques. The result of the simulated data of sample size 50 revealed that there is a difference between the two symmetric link functions with differing values of AIC with the Probit outperforming the logit link having least values of AIC which indicates that the probit link should be preferred under the symmetric assumption. While under the asymmetric link functions the loglog outperformed the cloglog with smaller values of AIC utilized on the life dataset which gives us the notion that the loglog link should be preferred under the asymmetric assumption. Furthermore table 6 also indicates that type of occupation is the only significant factor associated with hypertension in tuberculosis infected patients under study using both the cloglog and loglog link functions. On this note we recommend that patients with diabetes should be given less strenuous jobs and occupations to handle. Finally we were able to show that the link functions can be distinguished even with small values of (n < 1000) under the two assumptions.

  • Research Article
  • Cite Count Icon 16
  • 10.1265/ehpm.23-00271
Non-linear association between long-term air pollution exposure and risk of metabolic dysfunction-associated steatotic liver disease
  • Jan 1, 2024
  • Environmental Health and Preventive Medicine
  • Wei-Chun Cheng + 5 more

Metabolic Dysfunction-associated Steatotic Liver Disease (MASLD) has become a global epidemic, and air pollution has been identified as a potential risk factor. This study aims to investigate the non-linear relationship between ambient air pollution and MASLD prevalence. In this cross-sectional study, participants undergoing health checkups were assessed for three-year average air pollution exposure. MASLD diagnosis required hepatic steatosis with at least 1 out of 5 cardiometabolic criteria. A stepwise approach combining data visualization and regression modeling was used to determine the most appropriate link function between each of the six air pollutants and MASLD. A covariate-adjusted six-pollutant model was constructed accordingly. A total of 131,592 participants were included, with 40.6% met the criteria of MASLD. "Threshold link function," "interaction link function," and "restricted cubic spline (RCS) link functions" best-fitted associations between MASLD and PM2.5, PM10/CO, and O3 /SO2/NO2, respectively. In the six-pollutant model, significant positive associations were observed when pollutant concentrations were over: 34.64 µg/m3 for PM2.5, 57.93 µg/m3 for PM10, 56 µg/m3 for O3, below 643.6 µg/m3 for CO, and within 33 and 48 µg/m3 for NO2. The six-pollutant model using these best-fitted link functions demonstrated superior model fitting compared to exposure-categorized model or linear link function model assuming proportionality of odds. Non-linear associations were found between air pollutants and MASLD prevalence. PM2.5, PM10, O3, CO, and NO2 exhibited positive associations with MASLD in specific concentration ranges, highlighting the need to consider non-linear relationships in assessing the impact of air pollution on MASLD.

  • Research Article
  • Cite Count Icon 8
  • 10.1016/j.csda.2020.107013
Adaptive semiparametric estimation for single index models with jumps
  • May 23, 2020
  • Computational Statistics &amp; Data Analysis
  • Zhong-Cheng Han + 2 more

Adaptive semiparametric estimation for single index models with jumps

  • Research Article
  • Cite Count Icon 11
  • 10.1002/cpe.7005
On the performance of link functions in the beta ridge regression model: Simulation and application
  • Apr 10, 2022
  • Concurrency and Computation: Practice and Experience
  • Sidra Mustafa + 3 more

The beta regression model (BRM) is appropriate when the response variable is continuous and is in the form of ratios and proportions. For the estimation of the BRM, the maximum likelihood estimation (MLE) method is used with a specific link function. However, the MLE provides unstable results when the explanatory variables are correlated. In this study, we consider some ridge parameters for the beta ridge regression estimator (BRRE) under different link functions. However, mostly the researchers do not pay much attention to the suitable link function. So, we consider five link functions to see the performance of ridge parameters in the BRRE. For the performance assessment of ridge parameters and different link functions, a Monte Carlo simulation and a real application are considered, where mean squared error is used as the evaluation criterion. Both the simulation and example findings demonstrate that the BRRE with the log–log link function provides efficient results.

  • Dissertation
  • 10.5821/dissertation-2117-93822
Regression and residual analysis in linear models with interval censored data
  • Jul 19, 2002
  • Rebekka Topp

This work consists of two parts, both related with regression analysis for interval censored data. Interval censored data x have the property that their value cannot be observed exactly but only the respective interval [xL,xR] which contains the true value x with probability one.&lt;br/&gt;&lt;br/&gt;In the first part of this work I develop an estimation theory for the regression parameters of the linear model where both dependent and independent variables are interval censored. In doing so I use a semi-parametric maximum likelihood approach which determines the parameter estimates via maximization of the likelihood function of the data. Since the density function of the covariate is unknown due to interval censoring, the maximization problem is solved through an algorithm which frstly determines the unknown density function of the covariate and then maximizes the complete data likelihood function. The unknown covariate density is hereby determined nonparametrically through a modification of the approach of Turnbull (1976). The resulting parameter estimates are given under the assumption that the distribution of the model errors belong to the exponential familiy or are Weibull. In addition I extend my extimation theory to the case that the regression model includes both an interval censored and an uncensored covariate. Since the derivation of the theoretical statistical properties of the developed parameter estimates is rather complex, simulations were carried out to determine the quality of the estimates. As a result it can be seen that the estimated values for the regression parameters are always very close the real ones. Finally, some alternative estimation methods for this regression problem are discussed.&lt;br/&gt;&lt;br/&gt;In the second part of this work I develop a residual theory for the linear regression model where the covariate is interval censored, but the depending variable can be observed exactly. In this case the model errors appear to be interval censored, and so the residuals. This leads to the problem of not directly observable residuals which is solved in the following way: Since one assumption of the linear regression model is the N(0,&amp;#61555;2)-distribution of the model errors, it follows that the distribtuion of the interval censored errors is a truncated normal distribution, the truncation being determined by the observed model error intervals. Consequently, the distribution of the interval censored residuals is a -distribution, truncated in the respective residual interval, where the estimation of the residual variance is accomplished through the method of Gómez et al. (2002). In a simulation study I compare the behaviour of the so constructed residuals with those of Gómez et al. (2002) and a naïve type of resiudals which considers the middle of the residual interval as the observed residual. The results show that my residuals can be used for most of the simulated scenarios, wheras this is not the case for the other two types of residuals. Finally, my new residual theory is applied to a data set from a clinical study.

  • Book Chapter
  • Cite Count Icon 1
  • 10.1002/0470013192.bsa252
Generalized Linear Models (GLM)
  • Apr 15, 2005
  • Encyclopedia of Statistics in Behavioral Science

Generalized linear models provide a unified framework for regression models including multiple regression, logistic regression, analysis of variance, and analysis of covariance. Such models consist of three components – a link function specifying the transformation of the response variable to be modeled by a linear function of the explanatory variables, an error distribution suitable for each type of response, and a variance function specifying the relationship between the mean and variance of the error distribution.

  • Research Article
  • Cite Count Icon 9
  • 10.2307/2289307
A Nonparametric Approach to the Truncated Regression Problem
  • Sep 1, 1988
  • Journal of the American Statistical Association
  • Kwok-Leung Tsui + 2 more

A description is given of a new method of estimating the regression parameters in the linear regression model from data where the dependent variable is subject to truncation. The residual distribution is allowed to be unspecified. The method is iterative and involves estimation of the residual distribution under the truncated sampling scheme. The technique can be interpreted as an iterative bias adjustment of the observations in order to correct the regression relationship in the sampled population to match that of the model. A simulation study compares the performance of various estimators, including one suggested by Bhattacharya, Chernoff, and Yang (1983). This truncation regression problem arises in many contexts of scientific and social research. In economics Tobin (1958) analyzed household expenditure on durable goods using a regression model that took account of the fact that the expenditure is always nonnegative. A more general situation was studied by Hausman and Wise (1976, 1977) in connection with negative income-tax experiments. Another example concerning the schooling and earnings of low achievers was studied by Hansen, Weisbrod, and Scanlon (1970). There is also a controversy in astronomy involving Hubble's law and Segal's chronometric theory (Nicoll and Segal 1982; Turner 1979). Both theories predict a straight line relating the negative log of luminosity and the log of velocity as measured by red shift for celestial objects. The problem is complicated by the fact that objects of low luminosity are not visible, and hence all data relating to them are unobserved. Holgate (1965) described a biological example. A truncated linear regression model is defined as y = x T β + e, where x is a vector of covariates, β is the vector of parameter of interest, and e is independent of x with mean 0 and cumulative distribution F. The datum (x, y) is observed only if y ≤ y 0. The truncation point y 0 is known. Based on n independent observations (x i , y i ) with yi ≤ y 0, it is desired to estimate β and F. Note that this differs from the censored regression model where data (x, y) with y > y 0 is observed but with the y value set to y 0. The procedures described in the article are easily extended to truncation from below and the situation where the truncation points vary across observations. It is straightforward to see that the ordinary least squares estimate of β is inconsistent. A common method of dealing with this problem is to assume that the error distribution F is Gaussian and proceed with standard parametric methods. In many applications this assumption may not be reasonable. Hence there is interest in developing nonparametric methods of estimation that do not rely on assumptions about F. In this article a new approach for estimating β is introduced. The method allows the error distribution F to be arbitrary and is general enough to handle multiple linear regression. The rank-based method of Bhattacharya et al. (1983), which was designed for simple linear regression, is compared with the proposed method using a simulation study. The new approach appears to give estimators with good bias and efficiency properties in a wide variety of situations.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 4
  • 10.1007/s00180-023-01330-y
Bootstrapping binary GEV regressions for imbalanced datasets
  • Feb 4, 2023
  • Computational Statistics
  • Michele La Rocca + 2 more

This paper proposes and discusses a bootstrap scheme to make inferences when an imbalance in one of the levels of a binary variable affects both the dependent variable and some of the features. Specifically, the imbalance in the binary dependent variable is managed by adopting an asymmetric link function based on the quantile of the generalized extreme value (GEV) distribution, leading to a class of models called GEV regression. Within this framework, we propose using the fractional-random-weighted (FRW) bootstrap to obtain confidence intervals and implement a multiple testing procedure to identifying the set of relevant features. The main advantages of FRW bootstrap are as follows: (1) all observations belonging to the imbalanced class are always present in every bootstrap resample; (2) the bootstrap can be applied even when the complexity of the link function does not allow to easily compute second-order derivatives for the Hessian; (3) the bootstrap resampling scheme does not change whatever the link function is, and can be applied beyond the GEV link function used in this study. The performance of the FRW bootstrap in GEV regression modelling is evaluated using a detailed Monte Carlo simulation study, where the imbalance is present in the dependent variable and features. An application of the proposed methodology to a real dataset to analyze student churn in an Italian university is also discussed.

  • Research Article
  • 10.21754/iecos.v24i2.2003
Binary regression model with misclassification and Berkson-type measurement error with student-t distribution
  • Dec 31, 2023
  • revista IECOS
  • Marcos Antonio Alves Pereira + 1 more

In this article, we introduce a regression model tailored for fitting binary data affected by misclassification in the response variable and Berkson-type measurement error in the covariate. The conventional assumption of a normal distribution for measurement error may inadequately represent atypical observations present in the dataset. To address this limitation, our model incorporates misclassification in the response variable and Berksontype measurement error, employing the Student-t distribution for more robust modeling of these atypical observations. We utilize the cumulative distribution function from the Student-t distribution as the link function, enhancing our ability to capture the dataset’s unique characteristics. Model parameters are estimated via the maximum likelihood method. We conduct a comprehensive Monte Carlo simulation study to thoroughly assess the impact of measurement errors and misclassification. Additionally, we apply the proposed model to a real-world dataset of survivors from the atomic bombing in Japan, showcasing its adaptability and suitability in practical scenarios. Our findings highlight the robustness and flexibility of this model in effectively handling complex binary regression scenarios involving measurement errors and misclassification.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant