Skip to main content

Main menu

  • Home
  • Content
    • First Release
    • Current
    • Archives
    • Collections
    • Audiovisual Rheum
    • 50th Volume Reprints
  • Resources
    • Guide for Authors
    • Submit Manuscript
    • Payment
    • Reviewers
    • Advertisers
    • Classified Ads
    • Reprints and Translations
    • Permissions
    • Meetings
    • FAQ
    • Policies
  • Subscribers
    • Subscription Information
    • Purchase Subscription
    • Your Account
    • Terms and Conditions
  • About Us
    • About Us
    • Editorial Board
    • Letter from the Editor
    • Duncan A. Gordon Award
    • Privacy/GDPR Policy
    • Accessibility
  • Contact Us
  • JRheum Supplements
  • Services

User menu

  • My Cart
  • Log In

Search

  • Advanced search
The Journal of Rheumatology
  • JRheum Supplements
  • Services
  • My Cart
  • Log In
The Journal of Rheumatology

Advanced Search

  • Home
  • Content
    • First Release
    • Current
    • Archives
    • Collections
    • Audiovisual Rheum
    • 50th Volume Reprints
  • Resources
    • Guide for Authors
    • Submit Manuscript
    • Payment
    • Reviewers
    • Advertisers
    • Classified Ads
    • Reprints and Translations
    • Permissions
    • Meetings
    • FAQ
    • Policies
  • Subscribers
    • Subscription Information
    • Purchase Subscription
    • Your Account
    • Terms and Conditions
  • About Us
    • About Us
    • Editorial Board
    • Letter from the Editor
    • Duncan A. Gordon Award
    • Privacy/GDPR Policy
    • Accessibility
  • Contact Us
  • Follow Jrheum on BlueSky
  • Follow jrheum on Twitter
  • Visit jrheum on Facebook
  • Follow jrheum on LinkedIn
  • Follow jrheum on YouTube
  • Follow jrheum on Instagram
  • Follow jrheum on RSS
EditorialEditorial

Toward the Estimation of Unbiased Disease Prevalence Estimates Using Administrative Health Records

TITILOLA FALASINNU and JULIA F. SIMARD
The Journal of Rheumatology December 2019, 46 (12) 1549-1551; DOI: https://doi.org/10.3899/jrheum.190484
TITILOLA FALASINNU
Division of Epidemiology, Department of Health Research and Policy, Stanford School of Medicine;
PhD
Roles: Postdoctoral Fellow
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for TITILOLA FALASINNU
JULIA F. SIMARD
Division of Epidemiology, Department of Health Research and Policy, and Division of Immunology and Rheumatology, Department of Medicine, Stanford School of Medicine, Stanford, California, USA.
ScD
Roles: Assistant Professor
  • Find this author on Google Scholar
  • Find this author on PubMed
  • Search for this author on this site
  • ORCID record for JULIA F. SIMARD
  • For correspondence: jsimard{at}stanford.edu
  • Article
  • Figures & Data
  • Info & Metrics
  • References
  • PDF
PreviousNext
Loading

Data are information. And what we do with that information, how we process it, and interpret it can be complicated. It should not come as a surprise that these days there is a lot of talk about “big data” – about its promise, its potential, and its pitfalls. Big data (e.g., administrative, birth certificates, claims, electronic health records, registers) are growing in size, accessibility, and application. However, repurposing data from their original use to the research environment requires careful attention. Truthfully, whether we are talking about statistical analysis of small clinical datasets or supervised learning algorithms in big datasets, some of the same principles apply. No matter what, understanding where our data come from informs our design, our analysis, and most importantly, our interpretation.

There are 3 major sources of bias that determine whether inferences from a dataset are a close approximation of the truth: confounding, selection, and information. Confounding occurs when an association between 2 factors can be explained by an (often unmeasured) extraneous factor. Confounding often limits our ability to make truthful inferences about causality. Selection bias may occur when the choice of dataset limits the ability to generalize findings to the population affected by a disease. For example, using only drug claims data or hospitalization data to infer the prevalence of osteoarthritis (OA) may underestimate the condition because there may be individuals who may not need medication or have not been hospitalized in the time window evaluated. Information bias (often referred to as misclassification or measurement error) is also a threat to validity. Despite the potential problems of misclassification and measurement error, a recent systematic review found that fewer than 50% of studies from 12 high-impact journals in 2016 reported on this error, and only 7% used methods to assess or adjust for it1. Large samples alone cannot overcome systematic errors. In other words, infinitely large sample sizes will not necessarily mitigate these biases.

In this issue of The Journal, Slim, et al2, use health administrative data from Quebec to estimate the prevalence of rheumatoid arthritis (RA) in 2010. Using about 20,000 participants aged 40 to 69 years old from the large prospective CARTaGENE cohort, the authors linked to the Régie de l’assurance maladie du Québec for provincial administrative health data. The choice of this dataset for the estimation of the prevalence of RA is appropriate and limits selection bias because Canada has universal healthcare and the administrative datasets that house the physician billing data are a close approximation of the true patterns of disease burden, or at the very least what physicians diagnose and document in the electronic health records. The authors demonstrate how the data one chooses can influence results and also highlight the importance of addressing misclassification. By increasing the observation period, cases can be identified that may be milder, untreated, or managed predominantly by primary care during shorter time windows. Using Swedish population-based registry data from 2001 to 2007, we found a comparable prevalence of RA in 2008, presented age-stratified as 0.19% (40–49 yrs), 0.43% (50–59 yrs), and 0.89% (60–69 yrs) on the basis of visits to inpatient or outpatient specialist care or entry in the Swedish Rheumatology Quality Register3. The results are intuitive – adding self-reported data increased the prevalence compared to using administrative data alone, and adjusting for the potential false positives reduced the prevalence.

Across all modeling approaches to estimate the prevalence of RA, the authors showed how increasing the observation period of the data influences the estimated prevalence. With all followup time ending on December 31, 2010, the prevalence of RA increased as the duration of followup time increased. The authors and others have demonstrated this in other settings including those for systemic lupus erythematosus (SLE) and OA4,5,6,7,8. The authors also acknowledged that there are some pitfalls in the use of longer observation periods and self-reported information on RA, with both methods increasing the risk of the overestimation of the true prevalence. To combat these risks and boost the reliability of the prevalence estimates, they included misclassification error estimates and augmented self-reports of RA diagnosis with current use of disease-modifying anti-rheumatic drugs.

What this work adds to our dialogue is the importance of evaluating the likelihood of and accounting for potential misclassification. In the current setting, misclassification can happen 2 ways: (1) an individual is identified as having RA but does not actually have RA (a false positive), and (2) an individual is identified as not having RA but actually does (a false negative).

Remember that sensitivity is the probability of someone who has RA being identified by the algorithm as an RA case (true positive) and the specificity is the probability that an individual is correctly identified as not having RA (true negative). The complement of the latter, 1-specificity, is therefore the probability that an individual is falsely identified as a case [a false positive, i.e., Pr(Algorithm+|RA disease−)]. As the authors explain, the observed cases identified will always be some mix of true positives and false positives (i.e., a function of sensitivity and specificity). Hence, the authors use data on these variables informed by a published validation study and experts in the field, to adjust their estimates for this anticipated misclassification.

Misclassification and measurement errors are gaining more attention as sources of bias to be addressed in epidemiological research. Plotting the number of times these were mentioned by searching PubMed since 1995, we see that clearly more papers are considering this potential threat to validity (Figure 1). In the work by Slim, et al2, the authors applied Bayesian latent class analysis, incorporating prior sensitivity and specificity of the ascertainment methods into the likelihood function and acknowledging the unknown truth (i.e., gold standard of confirmed RA). Electronic health records–based phenotyping using Bayesian latent class analysis has also recently been applied to type 2 diabetes mellitus and is also covered elsewhere in detail9. Additional approaches, including non-Bayesian methods, and considerations for quantitative bias analysis are discussed by Lash and colleagues, including consideration of when bias analysis may be more helpful than necessary10.

PubMed search for papers that mentioned misclassification and measurement errors as sources of bias.
  • Download figure
  • Open in new tab
  • Download powerpoint
Figure 1.

PubMed search for papers that mentioned misclassification and measurement errors as sources of bias.

Confounding, selection bias, and misclassification are critical threats to the validity and generalizability of our work, and the size and availability of large datasets is only going to increase. In an extreme example, our group showed that death certificate data may underestimate the burden of SLE in Sweden, with about 59% of decedents with SLE lacked mention of SLE on their death certificates11. We determined that certain characteristics lead to missingness (older age and having a cancer diagnosis), and that there are situations where the extent and direction of the misclassification may be impossible to quantify. However, with administrative datasets such as the one used in this study, the level of misclassification is often not as stark as the limited death certificate data. We are getting more comfortable with the notion of confounding and the myriad strategies to tackle this potential source of bias. Measurement error and misclassification exist, whether we acknowledge them or not, and may not always simply lead to conservative estimates by biasing toward the null or diluting estimates of the truth. Thus, it is important to understand the provenance of the data and how that informs our interpretation.

Footnotes

  • See RA prevalence in Quebec, page 1570

REFERENCES

  1. 1.↵
    1. Brakenhoff TB,
    2. Mitroiu M,
    3. Keogh RH,
    4. Moons KGM,
    5. Groenwold RHH,
    6. van Smeden M
    . Measurement error is often neglected in medical literature: a systematic review. J Clin Epidemiol 2018;98:89–97.
    OpenUrl
  2. 2.↵
    1. Slim ZF,
    2. Soares de Moura C,
    3. Bernatsky S,
    4. Rahme E
    . Identifying rheumatoid arthritis cases within the Quebec health administrative database. J Rheumatol 2019;46:1570–6.
    OpenUrlAbstract/FREE Full Text
  3. 3.↵
    1. Neovius M,
    2. Simard JF,
    3. Askling J
    . Nationwide prevalence of rheumatoid arthritis and penetration of disease-modifying drugs in Sweden. Ann Rheum Dis 2011;70:624–9.
    OpenUrlAbstract/FREE Full Text
  4. 4.↵
    1. Ng R,
    2. Bernatsky S,
    3. Rahme E
    . Observation period effects on estimation of systemic lupus erythematosus incidence and prevalence in Quebec. J Rheumatol 2013;40:1334–6.
    OpenUrlAbstract/FREE Full Text
  5. 5.↵
    1. Nightingale AL,
    2. Farmer RD,
    3. de Vries CS
    . Systemic lupus erythematosus prevalence in the UK: methodological issues when using the General Practice Research Database to estimate frequency of chronic relapsing-remitting disease. Pharmacoepidemiol Drug Saf 2007;16:144–51.
    OpenUrlCrossRefPubMed
  6. 6.↵
    1. Wiréhn A-BE,
    2. Karlsson HM,
    3. Carstensen JM
    . Estimating disease prevalence using a population-based administrative healthcare database. Scand J Public Health 2007;35:424–31.
    OpenUrlPubMed
  7. 7.↵
    1. Powell KE,
    2. Diseker RA,
    3. Presley RJ,
    4. Tolsma D,
    5. Harris S,
    6. Mertz KJ,
    7. et al.
    Administrative data as a tool for arthritis surveillance: estimating prevalence and utilization of services. J Public Health Manag Pract 2019;9:291–8.
    OpenUrl
  8. 8.↵
    1. Kopec JA,
    2. Rahman MM,
    3. Berthelot J-M,
    4. Le Petit C,
    5. Aghajanian J,
    6. Sayre EC,
    7. et al.
    Descriptive epidemiology of osteoarthritis in British Columbia, Canada. J Rheumatol 2007;34:386–93.
    OpenUrlAbstract/FREE Full Text
  9. 9.↵
    1. Hubbard RA,
    2. Huang J,
    3. Harton J,
    4. Oganisian A,
    5. Choi G,
    6. Utidjian L,
    7. et al.
    A Bayesian latent class approach for EHR-based phenotyping. Stat Med 2019;38:74–87.
    OpenUrl
  10. 10.↵
    1. Lash TL,
    2. Fox MP,
    3. MacLehose RF,
    4. Maldonado G,
    5. McCandless LC,
    6. Greenland S
    . Good practices for quantitative bias analysis. Int J Epidemiol 2014;43:1969–85.
    OpenUrlCrossRefPubMed
  11. 11.↵
    1. Falasinnu T,
    2. Rossides M,
    3. Chaichian Y,
    4. Simard JF
    . Do death certificates underestimate the burden of rare diseases? The example of systemic lupus erythematosus mortality, Sweden, 2001–2013. Public Health Rep 2018;133:481–8.
    OpenUrl
PreviousNext
Back to top

In this issue

The Journal of Rheumatology
Vol. 46, Issue 12
1 Dec 2019
  • Table of Contents
  • Table of Contents (PDF)
  • Index by Author
  • Editorial Board (PDF)
Print
Download PDF
Article Alerts
Sign In to Email Alerts with your Email Address
Email Article

Thank you for your interest in spreading the word about The Journal of Rheumatology.

NOTE: We only request your email address so that the person you are recommending the page to knows that you wanted them to see it, and that it is not junk mail. We do not capture any email address.

Enter multiple addresses on separate lines or separate them with commas.
Toward the Estimation of Unbiased Disease Prevalence Estimates Using Administrative Health Records
(Your Name) has forwarded a page to you from The Journal of Rheumatology
(Your Name) thought you would like to see this page from the The Journal of Rheumatology web site.
CAPTCHA
This question is for testing whether or not you are a human visitor and to prevent automated spam submissions.
Citation Tools
Toward the Estimation of Unbiased Disease Prevalence Estimates Using Administrative Health Records
TITILOLA FALASINNU, JULIA F. SIMARD
The Journal of Rheumatology Dec 2019, 46 (12) 1549-1551; DOI: 10.3899/jrheum.190484

Citation Manager Formats

  • BibTeX
  • Bookends
  • EasyBib
  • EndNote (tagged)
  • EndNote 8 (xml)
  • Medlars
  • Mendeley
  • Papers
  • RefWorks Tagged
  • Ref Manager
  • RIS
  • Zotero

 Request Permissions

Share
Toward the Estimation of Unbiased Disease Prevalence Estimates Using Administrative Health Records
TITILOLA FALASINNU, JULIA F. SIMARD
The Journal of Rheumatology Dec 2019, 46 (12) 1549-1551; DOI: 10.3899/jrheum.190484
del.icio.us logo Twitter logo Facebook logo  logo Mendeley logo
  • Tweet Widget
  •  logo
Bookmark this article

Jump to section

  • Article
    • Footnotes
    • REFERENCES
  • Figures & Data
  • Info & Metrics
  • References
  • PDF

Related Articles

Cited By...

More in this TOC Section

  • From Mouth to Joint: Citrullinated Bacteria in Driving Synovial Autoimmunity
  • Bridging the Gap: Rheumatology Meets Palliative Care
  • More Than Skin Deep: Moving From a Skin-Based Phenotypic Model to Molecular Classification of Systemic Sclerosis
Show more Editorial

Similar Articles

Content

  • First Release
  • Current
  • Archives
  • Collections
  • Audiovisual Rheum
  • COVID-19 and Rheumatology

Resources

  • Guide for Authors
  • Submit Manuscript
  • Author Payment
  • Reviewers
  • Advertisers
  • Classified Ads
  • Reprints and Translations
  • Permissions
  • Meetings
  • FAQ
  • Policies

Subscribers

  • Subscription Information
  • Purchase Subscription
  • Your Account
  • Terms and Conditions

More

  • About Us
  • Contact Us
  • My Alerts
  • My Folders
  • Privacy/GDPR Policy
  • RSS Feeds
The Journal of Rheumatology
The content of this site is intended for health care professionals.
Copyright © 2025 by The Journal of Rheumatology Publishing Co. Ltd.
Print ISSN: 0315-162X; Online ISSN: 1499-2752
Powered by HighWire