• Skip to main content
  • Skip to secondary menu
  • Skip to primary sidebar
  • Skip to footer
  • Homepage
  • About Us
  • Advertising
  • Submission Instructions
  • Editorial Calendar
Amstat News

Amstat News

The Membership Magazine of the American Statistical Association

  • Printed Issues
  • Practical Significance Podcast
  • Additional Features
  • Columns
  • Member News
  • Departments
You are here: Home / Additional Features / Statistical Practice | The Diagnostic Said It Was Fine: Why the Standard Collinearity Check Passes Two Variables That Are Nearly the Same 

Statistical Practice | The Diagnostic Said It Was Fine: Why the Standard Collinearity Check Passes Two Variables That Are Nearly the Same 

October 1, 2026 Leave a Comment

Harshit Singh and Simoni Pant

Somewhere in your last analysis, there might be two predictors that are almost the same variable, and the collinearity check you ran to catch it told you everything was fine. It wasn’t lying. The calculation was right. The scale it reports on is the problem. Here’s how we found out. 

The Afternoon Everything Worked 

We were studying what smartphone features young buyers would pay for, and the first result was solid. Buyers wanted storage most and battery life least, and that order held everywhere we pushed it. Battery came last in every demographic subgroup we could form. 

We also had two hypotheses, both stated in advance. First, buyers spending their parents’ money would be more price sensitive than buyers spending their own. Second, younger buyers would be more sensitive than older ones. Tested separately, both came back significant and in the predicted direction. 

Then, we put both into one model—out of diligence rather than suspicion—and neither survived. Both p-values rose well above 0.05. The model itself stayed significant. 

Read that again. Something in the model was genuinely predicting price sensitivity. What it could not do was tell us which variable was responsible. Four of five of our younger respondents were parent-funded, and six of seven of the older ones paid themselves. It was nearly the same split of the sample, drawn twice, wearing different labels. 

Neither effect had vanished. Both still held on their own tests. What died was the ability to say which of the two was doing the work. 

Height by the Doctor, Height by the Trainer 

Imagine studying what predicts performance in basketball, and you enter each player’s height as measured by the team doctor and the same player’s height as measured by the trainer. Neither will be significant. Both will look useless. The model will predict well anyway, because the information is in there twice and the regression cannot decide which copy to credit. 

Nobody makes that mistake, because the redundancy is written into the names. Most redundancy is not labeled. 

The Number Cleared Every Check 

Our variance inflation factor was 1.58. Measured directly, the overlap between the two columns was 0.60. Those are the same fact written twice. For a binary pair, the variance inflation factor is a fixed function of the association between the columns, so one converts into the other exactly. The diagnostic did not miss the overlap. It reported it precisely—on a scale that didn’t set off any alarms. A variance inflation factor of 5, the threshold nearly every course teaches, corresponds to an association of 0.89, a hair short of the two variables being copies of one another. Everything below that clears the bar, including associations strong enough to make both coefficients uninterpretable. 

The scale is only half of it. A variance inflation factor tells you how much noisier your estimate has become. It doesn’t tell you the estimate is now answering a narrower question—not what the payment source is associated with, but what the payment source adds once age is accounted for. When two columns are nearly the same, that leftover is almost nothing. Both things sink the result. The diagnostic reports one. 

What We Would Do Differently 

Cross-tabulate your categorical predictors before you model, not after a result surprises you. One line of code. Had we run it first, the whole analysis would have been framed differently. Report the table itself, not only the statistic, so readers can see the overlap rather than take your word for it. 

Then, put the diagnostic on a scale your judgment can read. The variance inflation factor computes the right quantity and reports it on a scale calibrated for a different worry, and a reassuring scale is more dangerous than a loud one. So, convert it back: 

VIFAssociation Between the Two Predictors
1.50.58
20.71
2.50.77
50.89
100.95

For values not shown, R = √(1 – 1/VIF) 

This works, whatever your predictors are: R is a plain correlation for a continuous pair and Cramér’s V for a categorical one. With three or more predictors, it becomes one predictor against all the others together, not a single pair. 

Then ask the question no diagnostic will answer. Are these two things or one thing measured two ways? If you believe they are one thing, that belief usually implies something you can check. Ours did. Our payment variable had a middle category, cost shared with parents, and if dependence is a single dimension, sensitivity should rise steadily from self-funded through shared to parent-funded. An ordered-alternatives test said it does. That is a weaker claim than the two we started with—one the design can actually support. 

Decide in advance what you will do when you find it. The instinct is to drop one predictor. It is usually wrong, because the choice is arbitrary, and the surviving coefficient then carries both effects. Report the two as joint indicators of one construct, keep both pre-specified tests in, and make no claim about which variable is doing the work. 

What It Cost Us 

We had a mechanism we liked and a sharp prediction to go with it. Buyers spending someone else’s money face a justification burden, so they should resist hardest where the value is hardest to justify to the person paying. A storage upgrade is easy to defend. Four more hours of claimed battery life is not. The gap should have been widest on battery. 

We tested it and found nothing. The interaction was null, and no individual feature gap survived a correction for having tested five of them. The pattern there was ran backwards, with battery showing the smallest gap of the five. 

The Part That Took Longest to See 

Our own paper now reads as though we planned all of this. Its hypotheses section states the two variables are indicators of a single construct and we test their overlap directly rather than assume it away. Every word of that is true. We wrote it after the afternoon described above. 

That is what methods sections do. The surprise gets absorbed into the design, and what reaches print is a plan—which means that when this happens to you, you will not have read a single paper that showed it happening to anyone else. 

So, the paper is less satisfying than the one we thought we had. It is also what the data supports. The alternative was a tidier story, built on a variable we had accidentally counted twice. The reason those stories get published is not that researchers are dishonest. 

It’s that the diagnostic said everything was fine. 

It usually does. 

Editor’s Note: This article draws on research conducted by the two authors alongside Saloni Kumari, Keneithsulie Loucii, Yoshei Amanda Konyak, Sarika Kumari, Anushka Verma, and Harsh Rana at Shaheed Sukhdev College of Business Studies, University of Delhi. The full paper is currently under peer review; view a preprint with the detailed statistics and the underlying data. 

Harshit Singh

Harshit Singh is an honors research graduate in management studies from Shaheed Sukhdev College of Business Studies, University of Delhi. His research covers behavioral economics and comparative institutional analysis, and his writing on economics, strategy, and complex systems has been published by the Foundation for Economic Education, RealClearDefense, and other outlets.

    Simoni Pant

    Simoni Pant is a recent graduate in management studies from Shaheed Sukhdev College of Business Studies, University of Delhi. Her interests and work span marketing, business strategy, finance, and consumer insights, while her research focuses on inequalities in women’s autonomy, data-driven decision-making, and technology-enabled business solutions. 

      Filed Under: Additional Features, Statistical Practice Tagged With: collinearity check, diagnostics, Harshit Singh, predictors, Simoni Pant, Statistical Practice, University of Delhi, variance inflation factor

      Reader Interactions

      Leave a Reply Cancel reply

      Your email address will not be published. Required fields are marked *

      Primary Sidebar

      Search

      More to See

      2024–2025 Degree Data: Awarding of Statistics, Biostatistics Degrees Steady Amid Staggering Growth for Data Science, Analytics Degrees 

      October 1, 2026 By Olivia Brown

      Member Showcase: Robert Oster on the Power of Collaboration, Leadership, and Giving Back

      October 1, 2026 By Olivia Brown

      Who’s Missing from the Data? The Case for Disability Inclusion 

      October 1, 2026 By Olivia Brown

      Meet New Member Naren Prakash: Aspiring Statistician with a Love for Statistical Research 

      October 1, 2026 By Olivia Brown

      Data Science Certification

      ASA HOME

      American Statistical Association

      Communications from the Executive Director

      ASA Leader Hub

      ASA Career Connect

      STAFF LIST

      Kim Gilliam
      Naomi Friedman
      Amanda Malloy

      Archives

      Categories

      Footer

      Editorial Staff

      Managing Editor
      Megan Murphy

      Graphic Designers / Production Coordinators
      Olivia Brown
      Meg Ruyle

      Communications Strategist
      Val Nirala

      Advertising Manager
      Christina Bonner

      Contributing Staff Members
      Kim Gilliam

      American Statistical Association
      277 South Washington Street, Suite 370
      Alexandria, VA 22314-3646
      Phone: (703) 302-1857

       

      Copyright © 2026 · Magazine Pro on Genesis Framework · WordPress · Log in