Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> Some time later, a group of anonymous researchers downloaded those data, according to last week’s post on Data Colada. A simple look at the participants’ mileage distribution revealed something very suspicious. Other data sets of people’s driving distances show a bell curve, with some people driving a lot, a few very little, and most somewhere in the middle. In the 2012 study, there was an unusually equal spread: Roughly the same number of people drove every distance between 0 and 50,000 miles. “I was flabbergasted,” says the researcher who made the discovery.

It is kind of hilarious that they would fail at the simple task of generating fake but credible random data. Barely statistics literate? Probably using the simplest means available in Excel.



Honestly... it was probably an undergrad or contractor they tasked with fetching the data from the insurance company who decided to pad the numbers. Or maybe a desperate PhD student.

Which doesn't absolve the PI, who fundamentally has to stand behind the data, but it makes more sense than the PI being unable to fake convincing data.


What does PI stand for?


Principle Investigator. Basically, the individual (professor) responsible for a research project. They then employ a bunch of people to actually run the experiment.


I used to work as an insurance analyst and regularly dealt with customer-reported mileage distributions. It boggles the mind that this data was used for an academic paper when any junior analyst in an insurance context could tell you it's nonsensical with a glance. To me it just goes to show how far away these researchers are from the domain of the data that they're using in these studies. Kind of lowers the credibility of social sciences as a whole, unfortunately.


Well, the point of the study was too figure out who would be truthful, so it sounds like the data is exactly right then


There are more interesting studies to be done with this data IMO, which the researchers could have done if they had cared enough to talk to someone in the field.

As an example, we rarely saw completely normal "bell curves" with reported mileage. We often saw a roughly gaussian shape between 10k-30k, with a "J Curve" under 10k, where some % of dishonest people would report their mileage as absurdly low.

Where permitted by regulators, we would actually rate on this. If you had a single car, reported yourself as fully employed outside of the home, and also reported 5k mileage per year, you would receive a SURCHARGE compared to someone who reported 15k, because there was signal about your likelihood to make a claim in the fact that you were lying. The signal disappeared if you looked at people with more plausible arrangements, like having 2 cars, one of which had low mileage, had a single low-mileage vehicle but were self-employed (possibly WFH), etc...

I have to believe a clever researcher could find some interesting results with such data.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: