7 Important Data Science Papers

It is back-to-school time, and here are some papers to keep you busy this school year. All the papers are free. This list is far from exhaustive, but these are some important papers in data science and big data.

Google Search

PageRank – This is the paper that explains the algorithm behind Google search.

Hadoop

MapReduce – This paper explains a programming model for processing large datasets. In particular, it is the programming model used in hadoop.

Google File System – Part of hadoop is HDFS. HDFS is an open-source version of the distributed file system explained in this paper.

NoSQL

These are 2 of the papers that drove/started the NoSQL debate. Each paper describes a different type of storage system intended to be massively scabable.