Matei Zaharia

University of California, Berkeley

Papers

2

Total Citations

39

H-Index

2

About

Matei Zaharia is a pioneering computer scientist whose work has fundamentally reshaped how we process and analyze massive datasets. His primary research areas span distributed systems, cloud computing, and big data analytics, with a particular focus on making data-intensive computing accessible and efficient. Zaharia is best known as the creator of Apache Spark, a unified analytics engine that revolutionized big data processing by introducing in-memory cluster computing, achieving speeds up to 100x faster than traditional MapReduce for certain workloads. His seminal paper on Spark has accumulated over 5,000 citations, underscoring its transformative impact on both industry and academia. Beyond Spark, Zaharia contributed to the development of Apache Mesos, a cluster manager that enables efficient resource sharing across distributed frameworks. His work on large-scale estimation in cyberphysical systems, including arterial traffic estimation from sparse GPS data, demonstrates his ability to apply big data techniques to real-world challenges. As a co-founder of Databricks, Zaharia has successfully bridged the gap between cutting-edge research and practical deployment, making him a leading figure in the modern data ecosystem.

Research Focus

Key Achievements

2
H-Index
2
Papers
39
Total Citations
20
Avg Citations/Paper
🏆 Most Cited Paper
Large-Scale Estimation in Cyberphysical Systems Using Streaming Data: A Case Study With Arterial Traffic Estimation
36 citations · 2013
📈 Most Prolific Year: 2013 (1 Papers)
🤝 Key Collaborators: 4
🏛 Institutions: University of California, Berkeley

Top Papers

  1. 1
  2. 2

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 14 days ago