Does Timing Matter? Exploring the Effects of Measurement Error on Models
This study examines the impact of measurement error in independent variables, especially timing errors, on parameter inference in biological models. Synthetic data analyses show inference is generally robust, but oscillating systems are biased; the authors review statistical correction methods and emphasize assessing data quality to determine when such adjustments are necessary.
Measurement error is an unavoidable feature of experimental data collection. It is common in mathematical biology to consider measurement error in the dependent variable. However, less attention has been given to errors in the independent variable. This work is focussed on the effects of independent variable measurement error in the biological sciences and the available statistical methods to account for these errors when performing parameter inference. Through a series of synthetic data studies, the effects of various error models are investigated, with a particular focus given to error in the time a measurement is taken. Across many scenarios, parameter inference proves robust to these errors, even without directly accounting for them. However, we find some systems, such as oscillating systems, are particularly susceptible to these errors and parameter estimates become biased. To aid researchers in the biological sciences, we review some statistical methods to correct for measurement error. We assess the applicability of these methods in a biological context by considering data availability and necessary assumptions for the methods. We find measurement error can have non-trivial and counter-intuitive effects on parameter inference and suggest assessing the available data should be an integral step in the modelling workflow. This allows researchers to identify when the integration of statistical methods to correct for measurement error are warranted.
- Conference Article
2
- 10.1117/12.407630
- Nov 22, 2000
- Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE
Research into the near-infrared biomedical optical imaging has produced a multitude of inverse imaging algorithms. Recent experience has shown that when these algorithms are tested with experimental data, they falter due to a mismatch between observed and simulated measurements. When considering measurements for imaging, one must consider both measurement and model error. If data is recorded properly, then measurement error tends to be normally distributed with a mean of zero. Model error can be biased and spatially correlated due to inaccuracies in the diffusion approximation, inaccurate parameter estimates, numerical error, and other factors. This contribution discusses trends in the measurement and model error observed from measurements on a single-pixel, frequency domain photon migration system developed for biomedical optical imaging. In order to reduce the model error bias, an empirical approach was applied to find experimental variables that significantly affect it. This approach reduced the mean of the model error on a test data set and produced a slight smoothing effect on its distribution. Image reconstruction attempts show that the modified data set produces an improved image over the image reconstructed from the raw data set. To our knowledge, this is the first time that model and measurement error information have been incorporated into a three dimensional image reconstruction algorithm.
- Discussion
20
- 10.1086/687806
- Jul 1, 2016
- American Journal of Sociology
STILL SEARCHING FOR A TRUE RACE? REPLY TO KRAMER ET AL. AND ALBA ET AL.
- Research Article
- 10.25071/1874-6322.1298
- Jun 1, 2004
- Journal of Income Distribution®
Measurement error can have a significant impact on measures of inequality. Using a fairly flexible parametric specification of an independent multiplicative measurement error (IMME) model we explore the relationship between changes in the variance of measurement error, for a given mean of measurement error, on the Gini Coefficient. While the measured Gini is greater than the true Gini, the difference decreases as the variance of measurement error decreases. Copulas are used to relax the assumption of independence of measurement error and true income. In this case the measured Gini can be larger or smaller than the true Gini, depending on the correlation between true income and measurement error. Using the same approach with simulations the effect of a different distribution of measurement error is investigated.
- Research Article
7
- 10.1785/0120030153
- Jun 1, 2005
- Bulletin of the Seismological Society of America
The uncertainty in hypocenters and origin times depends on measurement error (the wrong onset is picked) and error in the travel-time tables (and thus earth model) used. The errors in the tables comprise a baseline shift (the average difference over all stations is not zero) and the residuals about the baseline. The residuals are usually referred to as model error. Baseline error only affects origin time and so is usually ignored. It is model error that can result in epicenter error. A priori variances of the model and measurement error are usually used to estimate the uncertainty on epicenters. Few estimates of these variances have been published. Here the two variances are estimated, relative to International Association for Seismology and the Physics of the Earth9s Interior (iaspei) 91, for the P times from explosions at the Nevada Test Site at stations at regional distances. The analysis shows that at a large signal-to-noise ratio (snr) the variance of the measurement error is 0.01 sec 2 . The measurement error increases as snr decreases. Further, the travel times appear to increase as snr decreases. Model error has a formal variance of up to 1.9 sec 2 , but this variance is irrelevant to assessing the uncertainty in epicenter estimates. It is systematic bias (if any) caused by model error that contributes to epicenter uncertainty. Without knowing the bias it is only possible to estimate the precision of an epicenter and this depends on the measurement error. The analysis gives estimates of model error for each source-to-station path, and these path effects can be used to correct the travel times to give a revised model. With correction for path effects the estimated uncertainty in the epicenters becomes a measure of the accuracy of a location. The results presented here show that when estimating these effects, care must be taken to ensure that variations in snr do not bias the estimates.
- Research Article
5
- 10.1111/sjos.12225
- Apr 6, 2016
- Scandinavian Journal of Statistics, Theory and Applications
Linear increments (LI) are used to analyse repeated outcome data with missing values. Previously, two LI methods have been proposed, one allowing non‐monotone missingness but not independent measurement error and one allowing independent measurement error but only monotone missingness. In both, it was suggested that the expected increment could depend on current outcome. We show that LI can allow non‐monotone missingness and either independent measurement error of unknown variance or dependence of expected increment on current outcome but not both. A popular alternative to LI is a multivariate normal model ignoring the missingness pattern. This gives consistent estimation when data are normally distributed and missing at random (MAR). We clarify the relation between MAR and the assumptions of LI and show that for continuous outcomes multivariate normal estimators are also consistent under (non‐MAR and non‐normal) assumptions not much stronger than those of LI. Moreover, when missingness is non‐monotone, they are typically more efficient.
- Research Article
- 10.1002/qj.5073
- Sep 24, 2025
- Quarterly Journal of the Royal Meteorological Society
The quantities of interest in geosciences (e.g., atmospheric wind, rock porosity) typically exhibit structure across a wide range of spatial and temporal scales. As a result, it is insufficient to define measurement errors as some deviations from a reference truth. The main issue is that different instruments and models filter reality in distinct ways. Without a precise mathematical description of this filtering, the exact nature of measurement, model, and representation errors remains unclear. To address this, a measurement model that specifies which geophysical field is being sampled and where is proposed. Based on this model, formal definitions for measurement, model, and representation errors are given. The distinction between these types of errors is also discussed. A novel aspect of this study is the demonstration of conditions under which measurement and representation errors will be correlated with one another. Using this new framework, two common measurement error estimation methods, namely the triple‐cornered hat and triple collocation, are shown to be special cases of a more general approach, called sampling‐aware field‐informed retrieval, that explicitly accounts for representation error. Idealized experiments are conducted to illustrate how this general approach performs in various measurement error estimation scenarios involving representation errors. In the more realistic scenarios where observations originate from differently sized sampling windows and are not perfectly collocated with one another, it is shown that estimating measurement errors is possible only if representation errors are accounted for. It is also demonstrated that displacement errors must be taken into account to avoid correlations between measurement and representation errors. More generally, this study advocates that the formulation of a measurement model is necessary for the measurement and representation errors to be clearly defined.
- Research Article
912
- 10.1002/uog.5256
- Feb 27, 2008
- Ultrasound in Obstetrics & Gynecology
Clinical practice involves measuring quantities for a variety of purposes, such as aiding diagnosis, predicting future patient outcomes, and serving as endpoints in studies or randomized trials. Measurements are almost always prone to various sorts of errors, which cause the measured value to differ from the true value; accordingly, studies investigating measurement error frequently appear in this and other journals. The importance of measurement error depends upon the context in which the measurements in question are to be used. For example, a certain degree of measurement error may be acceptable if measurements are to be used as an outcome in a comparative study such as a clinical trial, but the same measurement errors may be unacceptably large to make measurements usable in individual patient management, such as screening or risk prediction. In the past 20 years many papers have been published advocating how studies of measurement error should be analyzed, with a paper by Bland and Altman1 being one of the most cited and well known examples. There has been much controversy concerning the choice of parameter to be estimated and reported, and consequently confusion surrounding the meaning and interpretation of results from studies investigating measurement error. In this paper we first distinguish between the general concepts of agreement and reliability to aid researchers in considering which are relevant for their particular application. We then review the statistical methods that can be used to investigate and quantify agreement and reliability, dealing separately with the different types of measurement error study, while emphasizing the largely common techniques that should be used for data analysis. We reiterate that the judgment of whether agreement or reliability are acceptable must be related to the clinical application, and cannot be proven by a statistical test. We highlight the fact that reliability depends on the population in which measurements are made, and not just on the measurement errors of the measurement method. We discuss the advantages of method comparison studies making at least two measurements with each measurement method on each subject. A key advantage is that the cause of a correlation between paired differences and means in the so-called Bland–Altman plot can be determined, in contrast to when only a single measurement is made with each method. Throughout the paper, we try to emphasize that calculated values of agreement and reliability from measurement error studies are estimates of parameters, and as such we should report such estimates with CIs to indicate the uncertainty with which they have been estimated. We restrict our attention to measurements of a continuous quantity; alternative methods are required for categorical data2. One difficulty in the measurement error field is the number of different terms used to describe studies of measurement error. The terms 'agreement', 'reliability', 'reproducibility' and 'repeatability' are used with varying degrees of consistency in the medical literature. We first make clear the distinction between the statistical concepts of agreement and reliability3. Agreement quantifies how close two measurements made on the same subject are, and is measured on the same scale as the measurements themselves. Two measurements of the same subject may differ for a number of reasons, depending on the conditions under which the measurements were made. In a method comparison study there will be differences because of inherent variability in each of the measurement methods, as well as potentially a bias between the measurements from the methods. If the measurements are made by different observers or raters, differences may be due to biases between the observers. Agreement between measurements is a characteristic of the measurement method(s) involved, which does not depend on the population in which measurements are made, unless bias or measurement precision varies with the true value being measured. One popular way of quantifying agreement is to estimate the 95% limits of agreement, as proposed by Bland and Altman1. These limits are defined such that we expect that, in the long run, 95% of future differences between measurements made on the same subject will lie within the limits. If reliability is high, measurement errors are small in comparison to the true differences between subjects, so that subjects can be relatively well distinguished (in terms of the quantity being measured) on the basis of the error-prone measurements. Conversely, if measurement errors tend to be large compared with the true differences between subjects, reliability will be low because differences between measurements of two subjects could be due purely to error rather than to a genuine difference in their true values. The reliability parameter is also known as an intraclass correlation (ICC), as it equals the correlation between any two measurements made on the same subject. Reliability takes values between zero and one, with a value of one corresponding to zero measurement error and a value of zero meaning that all the variability in measurements is due to measurement error. As a dimensionless quantity, it is arguably quite difficult to interpret, and deciding what value constitutes sufficiently high reliability is often made in a subjective fashion. Repeatability of measurements refers to the variation in repeat measurements made on the same subject under identical conditions4. This means that measurements are made by the same instrument or method, the same observer (or rater) if human input is required, and that the measurements are made over a short period of time, over which the underlying value can be considered to be constant. Variability in measurements made on the same subject in a repeatability study can then be ascribed only to errors due to the measurement process itself. Reproducibility refers to the variation in measurements made on a subject under changing conditions4. The changing conditions may be due to different measurement methods or instruments being used, measurements being made by different observers or raters, or measurements being made over a period of time, within which the 'error-free' level of the variable could undergo non-negligible change. The first type of study we consider is a repeatability study, in which we investigate and quantify the repeatability of measurements made by a single instrument or method, and in which the conditions of measurement remain constant. The second type of study we consider is a method comparison study, in which measurements are made using two measurement methods on a sample of subjects. The use of two different methods means that this is a reproducibility study, but the term method comparison is used as it clearly communicates the changing conditions under which measurements have been made. In contrast to a repeatability study, systematic bias may exist between measurements made by two different methods, and their measurement errors may have different SDs. The last type of study we consider is one in which measurements are made by different observers or raters. Again this is a type of reproducibility study. As with method comparison studies, biases may exist between observers, and their measurement SDs may differ. As we discuss below (see Measurements with observers or raters), if interest lies in quantifying the measurement error characteristics of the particular observers in one's study, exactly the same analysis methods should be used as for a method comparison study. In order to investigate the repeatability of measurements, a repeatability study must, for an appropriately selected sample, make at least two measurements per subject under identical conditions. This means that the measurements must be made by the same measurement method, or the same observer or rater. The objective is to then quantify the agreement and reliability of measurements made by that particular method or observer. If the differences between two measurements made on a subject are approximately Normally distributed, in the long run we expect the absolute difference between two measurements on a subject to differ by no more than the repeatability coefficient on 95% of occasions. To estimate the within-subject SD, we can fit a one-way analysis of variance (ANOVA) model to the data containing the repeat measurements made on subjects. ANOVAs can be fitted in all modern statistical packages. ANOVA models partition variability in data into that which can be ascribed to differences between groups, and that remaining to within groups. For a repeatability study, the groups are defined by the subjects under measurement, and this must be specified in the statistical package used. The ANOVA model estimates how much variation in measurements can be attributed to differences in the true, or 'error-free' values of subjects, with the remainder constituting measurement error. Fitting the ANOVA model results in estimates of the between-subject and within-subject SDs (or alternatively the corresponding variances, which are the SDs squared). The estimate of the within-subject SD can be used in the above formula to give an estimate of the repeatability coefficient. We illustrate with a study by Järvelä et al., who investigated the repeatability of measurements of flow volume made by one observer using three-dimensional (3D) power Doppler ultrasonography from 22 ovaries5. The estimated repeatability coefficient was 3.97, meaning that the absolute difference between any two future measurements made by that particular observer on a particular subject/unit are estimated to be no greater than 3.97 on 95% of occasions. It is important to note that the repeatability of another observer may be different, because of differences in the training and ability of observers. Because the repeatability coefficient calculated is an estimate, it is important to calculate a CI for it to indicate how precisely it has been estimated. A CI may be given automatically by statistical software, but in the Appendix we review how a 95% CI for the within-subject SD can be calculated. Such a CI can be used to find a CI for the repeatability coefficient by multiplying the CI limits by . If the CI for the within-subject variance is given by software instead, the limits must first be square rooted to give a CI for the within-subject SD. The ANOVA model assumes that the measurement errors are statistically independent of the true 'error-free' value, and that the SD of the errors is constant throughout the range of 'error-free' values. Sometimes the SD of errors increases with the true value being measured. This should be checked by plotting paired differences between measurements against their mean, the so called Bland–Altman plot. We illustrate this later in the context of a method comparison study (see Plotting the data) and describe how such heterogeneous errors can often be dealt with by making a log transformation (see Non-uniform variability in measurement of the repeatability coefficient on the differences between measurements being approximately Normally This can be checked by a or plot of the paired differences in measurements on each subject. made in the of the CI for the within-subject SD is that the measurement errors are If this is in the may be to The reliability of a measurement method is often of interest when measurements are to be used to between subjects or groups of subjects. For example, if we have a choice of two measurement methods that could be used to an outcome in a clinical or study, using the method with reliability will give greater statistical power to differences between groups, for a given sample In this we describe how measurements from a single method can be used to estimate the reliability in a given we discuss the use of reliability to measurement methods (see Reliability in method comparison and different observers (see in The same ANOVA model as above (see can be used to estimate statistical will automatically give the estimated which is the of we can the estimates of the and within-subject SDs into the formula given above (see Agreement and to estimate the For the flow data of Järvelä et an estimated of was This means that of the variability in measurements of the flow was estimated to be due to genuine differences in flow between with the remaining being due to errors in the measurement process and the observer Because the measurements were all made by one the reliability may be to as As with the repeatability the calculated is an estimate, and a 95% CI should be statistical give a but in the Appendix we review the of a 95% CI for the For the flow data of Järvelä et the estimate of a 95% of to Because we use the same ANOVA model to estimate reliability as we describe for agreement, the same of 95% CIs for the estimate on of the true 'error-free' values. The reliability of a measurement method depends upon the of the population in which the measurements are made. the of reliability given above (see Agreement and we that the of true 'error-free' values in the measured by the between-subject SD, the value of agreement between repeat measurements is a characteristic of the method or instrument the of measurement errors is the range of true reliability depends on the of measurement errors and the true in the population in which measurements are made. The is with a a method for measuring has a within-subject SD of which does not with the underlying value being measured. we a study to estimate the reliability, and we sample from a heterogeneous population in which the between-subject SD in true is 20 This give a reliability or of . we sample our subjects from a more in which the SD of true values is the same as the variability of the within-subject error. In this the reliability is to . If studies report only an estimate of the reliability (ICC), can only make use of the estimate if the population in which the to use such measurements has We that report estimates of and within-subject SD, in to the In this can whether the measurement method will be sufficiently for their application, in which the between subjects may be we use a measurement method or in clinical we must that the measurements it are to by the measurement method used, that measurements made with the method are using the method. If the measurements from the two methods are sufficiently close and the of a patient on their basis be the the method could the method in clinical because the method is or to The as to what constitutes depends on how the measurements are to be used. If the two measurements are made on the same or we can quantify their To investigate and quantify the agreement between measurements made by two methods, we must at a a sample of subjects using the measurement method and the proposed method. The data from such a study of of measurements from each with the containing the measurements from the two measurement methods. we making two measurements on each subject using each method because the one measurement per method is the most common we by analysis. The first to such a is to plot the The plot is of measurements from the method against from the method (or If measurements were from we expect the to lie on the of the of can be used to if there is if more data are above the than this that the method on the measurements on the of fit for this plot and report a statistical for whether the of the plot from the of As by Bland and in a paper in this we expect the to have a than one if the method on the any measurement and so a of the that the true is to one is not a it the same as a plot with the of of the between the measurements from two methods is often more by plotting the difference in a measurements from the two methods against the of their measurements, as first by and this is to as the Bland–Altman and is frequently in of measurement error We illustrate the plot using an from the literature. et volume measurement in with and volume was measured using and also by We have approximately their plot of differences against means and it is as using only data from the in volume measured by three-dimensional and by against their from study by et the SD and the SD, each with 95% CI The plot the difference between calculated using and against the of two values. et not their plot or of results to indicate which measurements were from It is important to be and always the same measurement from the The the of the paired from zero an estimate of the bias between the two methods. which measurements have been from which to which method, on measurements. The indicate the estimated limits of agreement their CI which we describe The plot can be used to the differences between measurements made by the two methods. The variability of the differences between the two methods how well the methods If the variability of the paired differences is the range of measurements, and there is no between the difference and mean, we can quantify the agreement between the methods by the limits of the of the limits of agreement, we illustrate how the for limits of agreement may be and how they can be dealt on we expect the difference in volume as measured by the two methods to lie between and for 95% of future measurements. It is important to that the limits of agreement calculated in this are just and as with any type of estimate it is to quantify and report how precisely the limits are estimated of 95% The estimated limits of agreement indicate how large the between the measurements will be on 95% of occasions. The whether the methods sufficiently well must then be made depending on the context in which the measurements will be used. In contrast to the repeatability which assumes no bias between measurements, the limits of agreement method this The of the paired differences whether on one method to or measurements to the measurements of the second method, which we to as a bias between the methods. In the data in from et al., there was of a bias between the two methods, because the difference between the methods was close to at We can a statistical to whether there is of bias by a of the differences paired of the measurements from each we the that the true of the differences is corresponding to no bias between the methods. For the data in this that there was no of a bias between and made using The limits of agreement method assumes that the SD of the differences is throughout the range of measurements. Sometimes this will not be the In frequently the SD of the differences will with the In this for small values the limits of agreement will be than for values the limits will be of the plot of differences against from the study by et that the variability of the differences may be when the value being measured is The SD of the paired differences below the of the means was it was for above the This can often be by of the measurements of the two If the difference in the of the measurements results in constant we can calculate the limits of agreement and CIs for the limits in the Because the difference of two is to the of the we can all the estimates made on the log scale by Because of the limits of agreement to the of one measurements to the For the data in the of the difference in measurements was with SD The limits of agreement on the log scale are to the of with 95% limits of agreement for the of to The 95% CIs can be calculated on the log scale to the which measurements were made by which method, we cannot whether the is for measurements by to by or a plot of the of the measurements by the two methods to their mean, using a with the estimated 95% limits of agreement and their 95% This that the of a SD for the paired differences is more of measured by three-dimensional (3D) and by against of measurements, on data published by et the SD and the SD, each with 95% CI to log transformation can also be used to try to For example, et estimate limits of agreement using paired differences in measurements as a of the of the two measurements. The key is that a transformation should be used which that are approximately Normally and have approximately SD the range of It is for there to be an between the paired differences and of this is in a study by et al., who the agreement between as measured using and the calculated using in We have the data from their Bland–Altman plot of the difference between the volume and volume against the mean, and it in There to be a between the paired differences and differences between volume measured by three-dimensional and by from study by et correlation the SD and the SD, each with 95% CI We can a statistical to the for a whether the correlation coefficient between the paired differences and means from zero or by of the differences against the The is known as in the context of the SD of measurements from two methods, which for the data in is statistically for an between paired differences and how should we than models for the difference as a of the we should be given to the cause of the A between the paired differences and means when the SD of measurements from one method from the SD of measurements from the other method. There are at least two of an between paired differences and The first is that there is between the difference in measurements from the two methods and true value being the bias between methods over the range of true values. The second is that the within-subject SDs of the two methods differ. This will in the of changing bias if a method has or measurement errors than the method. This is and when a or measurement method is compared with a If the cause of the correlation is within-subject SDs and we were to plot the paired differences against the true value being there be no The between the Bland–Altman plot and a plot against true values because the paired means measurement errors, in contrast to the true values. as by and with only one measurement per method per we cannot which of two is the To we must make an that the within-subject SDs of the two methods are so that the correlation is attributed to changing or that the two within-subject SDs are and that there in no between bias and true value we may have a of For the of we with the from et by that the correlation is due to a difference in the within-subject SDs of the two methods. A correlation between differences and means will when the method with measurement error is from measurements of the method with error. For the data in the correlation is the errors than the measured by which and to the volume measured by as the true that such measurements were to no an that there is no changing we can calculate 95% limits of agreement, as by et al., because the limits of agreement method does not that the two within-subject SDs are In order to from the data whether changing bias is we must make at least two measurements with each method on each subject. This study has often been but is we describe how such studies can be to whether a correlation between paired differences and means is due to changing bias between the methods (see with two measurements from each As reliability may be a parameter with which to two different measurement methods. To estimate each reliability, we must make at least two measurements of each subject with each of the two methods. The repeat measurements from each method can then be as two repeatability studies (see estimates of each reliability, which can be advantage of using reliability to measurement methods is that it can be used to methods when their measurements are given on different or as the reliability is a dimensionless Because reliability depends on the of the true values in the population (see Reliability depends on population it is that reliability are compared only if they have been estimated from the same report a single reliability estimate from method comparison studies in which it that only a single measurement per subject per method is It appear that the estimates are using a one-way ANOVA as above (see The one-way ANOVA model assumes that there are no systematic biases between measurements within subjects, and that the within-subject SDs are for all measurements measurements are from two different methods, of may be bias may exist between the methods and their within-subject SDs may differ. If is the model is and the estimates of within and between subject SD are estimates of different The reliability estimate then has a different interpretation from the We against the estimate from the one-way ANOVA to data from two different measurement methods. As making two measurements with each method an of whether bias between the methods is constant or whether the measurement error SDs differ. the measurements from each method can be separately as two repeatability studies, using the methods estimates of the repeatability coefficient and reliability for each method. The Bland–Altman plot of paired differences against means can be by using the of a two measurements from each method in of the single As with the Bland–Altman plot on a single measurement from each method, an correlation between differences and means could be due to changing bias or a difference in measurement error between the methods. We describe a to whether any bias between measurements made by a method with the as measured by the method. We by the of the two measurements made using the method and we by the measurement made on subject by the measurement method. A plot of against can then be used to investigate whether the bias between the and methods with the as measured by the method. It is that is against rather than against because the a to the common measurement error in and It is for the same that plotting paired differences between measurements from two methods against the measurements from one method is in method comparison between values of and that the bias between the methods with the true value as measured by the method. A statistical of this can be by a for with as the The for the that the in this from zero the of against the of constant If of bias changing with true value is we may be in the at which the bias between the methods We can estimate this by the estimate from the above model by our estimate of the This quantity is an estimate of the in bias method for a in the true value as measured by the method. The choice to against is We could have against and have a different and estimate of the at which the bias This choice is because we will two different and because it means the is statistically The model underlying is a type of and as such can be fitted using software for The is to and for measurements are made by observers or raters. measurements made by two different observers are than are two measurements made by the same observer. as with two methods, measurements from two observers may differ due to bias between the observers observer and their measurement errors may also have different SDs. For example, measurements from an observer who can make more measurements will have a SD than made by a observer. If measurements in the future are to be made by different observers, we to describe and quantify the differences between such measurements in order to whether differences are genuine or may be due to measurement error. As with a method comparison study, the way to study this is for each observer to make at least two measurements of a sample of subjects. The of such a study and the type of statistical analysis that are should be by whether interest lies in a particular of observers, or whether we are in a population of observers or raters. If future measurements are to be made by a of observers, the measurement error study should each of observers making measurements of a sample of subjects, with at least two measurements per observer per subject. We can then use exactly the same methods as for method different observers as different measurement methods. If each observer two or more measurements on each we can whether there is bias or changing bias between observers. We can the bias of future measurements (in by measurements using the corresponding estimated We can also estimate repeatability and reliability for each using the one-way ANOVA model (see It may be of interest to which observers are more and if differences in reliability can be related to observer such as of or If we are to that biases between observers are we can fit a so-called model to such a for a subject and observer The estimates indicate the and for bias between observers. such models can be to the measurement error SD to differ between observers. The observers in a measurement error study can often be considered as a sample of observers from a population of observers who may be used in future studies or clinical In this we are not in the particular observers in the measurement error study, but only in the that they the population of observers. In this it is important that a number of observers is used in the measurement error study. For example, if the measurement error study only two observers, can be the population of observers, because we have a sample of just Because the observers in the measurement error study are considered a sample from the population of observers, we such a study using a model that the observer as a We use a with subject and observer Fitting a model estimates of the between-subject SD, SD and the measurement error SD to be the same for different The SD variability in measurements due to observers, or biases between the observers. such estimates one can calculate the SD of the difference in two measurements, made by the same observer or by two different observers. the SD is the will be than the This is because biases between observers to make measurements from different observers it is more difficult to distinguish between subjects on the basis of measurements made by two different observers than if the subjects been measured by the same observer. In this paper we have distinguished between the concepts of agreement and parameter is to the as they describe different characteristics of the measurement The choice of what to report in a particular study should be by how measurements are to be used in the and also by the fact that may to use a measurement method in a different We have the fact that the reliability of a method depends not only on the of the measurement errors but on the of true values in the population in which measurements are made. As measurement techniques potentially may be used in a variety of clinical and different it is to report estimates of and within-subject SDs. We have which methods we are for the analysis of repeatability studies, method comparison studies, and studies with measurements made by different observers or raters. In we not a single reliability coefficient should be used for method comparison If the reliability of two methods are to be each reliability should be estimated by making at least two measurements on each subject with each measurement method. We have how an between paired differences and means may not be by changing bias between two methods. Such an may also be by a difference in the measurement error but with only one measurement per subject per method it is not to which of is the comparison studies should make at least two measurements per subject per method. This an of the of any between paired differences and means for measurements made by the two methods, and also the repeatability and reliability of each method to be estimated. measurements an observer or measurement error studies must use an number of observers if interest lies in making a population of observers. we results for the of CIs for two in the one-way ANOVA in which two measurements are from each
- Research Article
- 10.1007/s10608-026-10719-0
- Mar 16, 2026
- Cognitive Therapy and Research
Background To better understand the development of mental disorders, dynamic networks have gained more attention in recent years. Most of these network models use a single indicator per node despite the fact that measurement error may bias parameter estimates. This can lead to incorrect conclusions about the presence or absence of edges, as well as the relative strength of edges in the network. In this study, we compared single-indicator dynamic networks to approaches using information on multiple indicators per node to account for measurement error. Data and Methods We conducted two simulation studies, using time series (Study 1, N = 1) and panel (Study 2, N > 1) data, to compare the estimation of network parameters in the presence of measurement error in models with single indicators versus models with multiple indicators, namely as latent variables, plausible values, factor scores, and average scores. Across conditions, we varied the variance of the measurement error and the number of observations (in time series: number of timepoints; and in panel data: number of persons and waves). We evaluated the performance of each model by examining the correlation between the estimated and true network edge weights, as well as the sensitivity, specificity, and precision. Results In both studies, measurement error decreased correlations between the true and estimated network as well as sensitivity among all approaches, while specificity and precision were mostly unaffected. The single-indicator approach was the most sensitive to measurement error and the number of observations compared to other approaches. In Study 1, the factor and average score approaches performed best for temporal networks, and the latent variable approach for contemporaneous networks. In Study 2, generally the best-performing approach was the plausible value score. Discussion Measurement error may substantially bias estimates in dynamic networks, and multiple-indicator approaches can mitigate this bias. Multiple-indicator approaches generally outperformed the single-indicator approach, but the choice between different multiple-indicator approaches depends on several factors that must be carefully considered before deciding the best method for each study.
- Conference Article
- 10.1063/5.0126025
- Jan 1, 2023
- AIP conference proceedings
Artificial Intelligence Algorithms have been used in recent years in many scientific fields. We suggest employing artificial TABU algorithm to find the best estimate of the semi-parametric regression function with measurement errors in the explanatory variables and the dependent variable, where measurement errors appear frequently in fields such as sport, chemistry, biological sciences, medicine, and epidemiological studies, rather than an exact measurement.We estimate the regression function of the semi-parametric model by estimating the parametric model and estimating the non-parametric model, the parametric model is estimated by using an instrumental variables method (Wald method, Bartlett's method, and Durbin's method), The non-parametric model is estimated by using kernel smoothing (Nadaraya Watson), K-Nearest Neighbor smoothing and Median smoothing. The TABU algorithms were employed and structured estimating the semi-parametric regression function with measurement errors in the explanatory and dependent variables, then compare the models to choose the best mode, where the comparison between the models is done using the mean square error (MSE).A simulation had been used to study the empirical behavior for the semi-parametric models, with different sample sizes and variances. The most important conclusions that we reached when using statistical methods in estimating parameters and choosing the best model, we found that the Median-Durbin model is the best as it has less MSE, but when using Tabu algorithm showed that the Median-Wald model is the best because it has the lowest MSE.
- Research Article
2
- 10.22075/ijnaa.2022.5744
- Jan 1, 2022
- International Journal of Nonlinear Analysis and Applications
Artificial Intelligence Algorithms have been used in recent years in many scientific fields. We suggest employing flower pollination algorithm in the environmental field to find the best estimate of the semi-parametric regression function with measurement errors in the explanatory variables and the dependent variable, where measurement errors appear frequently in fields such as chemistry, biological sciences, medicine, and epidemiological studies, rather than an exact measurement. We estimate the regression function of the semi-parametric model by estimating the parametric model and estimating the non-parametric model, the parametric model is estimated by using an instrumental variables method (Wald method, Bartlett's method, and Durbin's method), The non-parametric model is estimated by using kernel smoothing (Nadaraya Watson), K-Nearest Neighbor smoothing and Median smoothing. The Flower Pollination algorithms were employed and structured in building the ecological model and estimating the semi-parametric regression function with measurement errors in the explanatory and dependent variables, then compare the models to choose the best model used in the environmental scope measurement errors, where the comparison between the models is done using the mean square error (MSE). These methods were applied to real data on environmental pollution/ air pollution in the city of Baghdad, and the most important conclusions that we reached when using statistical methods in estimating parameters and choosing the best model, we found that the Median-Durbin model is the best as it has less MSE, but when using flower The pollination algorithm showed that the Median-Wald model is the best because it has the lowest MSE, and when we compare the statistical methods with the FPA in selecting semi-parametric models, we notice the superiority of the FP algorithm in all methods and for all models.
- Research Article
30
- 10.1016/j.ymssp.2018.05.007
- May 11, 2018
- Mechanical Systems and Signal Processing
On choice and effect of weight matrix for response sensitivity-based damage identification with measurement and model errors
- Research Article
54
- 10.1111/j.2517-6161.1990.tb01791.x
- Jan 1, 1990
- Journal of the Royal Statistical Society Series B: Statistical Methodology
SUMMARY Hypothesis tests in generalized linear models are studied under the condition that a surrogate w is observed in place of the true predictor x. The efficient score test for the hypothesis of no association depends on the conditional expectation E(x|w) which is generally unknown. The usual test substitutes w for E(x|w) and is asymptotically valid but not efficient. We investigate two new test statistics appropriate when w = x + z where z is an independent measurement error. The first is a Wald test based on estimators corrected for measurement error. Despite the correction for attenuation in the estimator, this test has the same local power as the usual test. The second test employs an estimator of E(x|w) and is both asymptotically efficient for normal errors and approximately efficient when the measurement error variance is small.
- Research Article
2
- 10.3390/w14040522
- Feb 10, 2022
- Water
The hydraulic parameters representative of actual aquifer conditions can be obtained through aquifer tests formerly known as pumping tests. Diverse methodologies based on analytical or numerical solutions have been proposed for the interpretation of aquifer tests; however, measurement and model errors are often neglected, which could lead to hydraulic parameter values that do not reflect the aquifer conditions. In this paper, a new alternative is presented for the interpretation of aquifer tests in confined aquifers based on the Cooper–Jacob solution by means of the dynamic Kalman filter and a nonlinear optimization method. This proposal was tested in two previously published case studies; the measured drawdowns were filtered by considering measurement and model errors to match the Cooper–Jacob solution. For the case studies, the results show that filtering the measured drawdowns leads to variations of up to 49.97% in the values for T and 150% for S when compared to the values determined by methodologies that neglect measurement and model errors. A poor match between filtered and measured data reflects large measurement errors and considerable deviations of the aquifer conditions with respect to the proposed model.
- Research Article
15
- 10.1139/cjfr-2015-0265
- May 1, 2016
- Canadian Journal of Forest Research
In estimating aboveground forest biomass (AGB), three sources of error that interact and propagate include (i) measurement error, the quality of the tree-level measurement data used as inputs for the individual-tree equations; (ii) model error, the uncertainty about the equations of the individual trees; and (iii) sampling error, the uncertainty due to having obtained a probabilistic or purposive sample, rather than a census, of the trees on a given area of forest land. Monte Carlo simulations were used to examine measurement, model, and sampling errors and to compare total uncertainty between models and between a phase-based terrestrial laser scanner (TLS) and traditional forest inventory instruments. Input variables for the equations were diameter at breast height, total tree height (defined the height from the uphill side of the tree to the tree top), and height to crown base; these were extracted from the terrestrial LiDAR data. Relative contributions for measurement, model, and sampling errors were 5%, 70%, and 25%, respectively, when using TLS, and 11%, 66%, and 23%, respectively, when using the traditional inventory measurements as inputs into the models. We conclude that the use of TLS can reduce measurement errors of AGB compared with traditional inventory measurements.
- Research Article
4
- 10.1080/02331880108802726
- Jan 1, 2001
- Statistics
On a repeated-measurement model with errors in dependent variable