2 hours ago · Tech · hide · 0 comments

I'm at JSM, which is always a fantastically rich event. One is awed by the variety and breadth of topics over hundreds of sessions of talks. Each year, someone will mention a pet peeve. The other day, a speaker explains why "practitioners" often decide to drop missing data instead of imputing them. By his observation, most real-world datasets have tons of missing values. Many variables may be up to 40% missing. Given this practical fact, he thinks it obscene to impute that amount of missing data. You'd be making a lot of assumptions.The flip side of the argument is that when you don't impute, and simply drop the units with missing data, you do not make any assumptions, and most importantly you do not "fabricate" any data.Sure, imputation requires making some assumptions, like any statistical procedure does.But dropping missing data also makes one assumption.And this one assumption is usually the worst possible assumption one could make!The assumption is that the units with missing…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.