Data Science: Capstone
15-20 hours a week • Start today
Individual Course
Course Length
8 weeks
2-4 hours a week
Featuring faculty from:
Harvard T.H. Chan School of Public Health
Enroll as Individual
Certificate Price:
$ 149
Enroll as Individual
Certificate Price:
$ 149
In this online course taught by Harvard Professor Rafael Irizarry, build a movie recommendation system and learn the science behind one of the most popular and successful data science techniques.
Perhaps the most popular data science methodologies come from machine learning. What distinguishes machine learning from other computer guided decision processes is that it builds prediction algorithms using data. Some of the most popular products that use machine learning include the handwriting readers implemented by the postal service, speech recognition, movie recommendation systems, and spam detectors.
In this course,part of our Professional Certificate Program in Data Science, you will learn popular machine learning algorithms, principal component analysis, and regularization by building a movie recommendation system.
You will learn about training data, and how to use a set of data to discover potentially predictive relationships. As you build the movie recommendation system, you will learn how to train algorithms using training data so you can predict the outcome for future datasets. You will also learn about overtraining and techniques to avoid it such as cross-validation. All of these skills are fundamental to machine learning.
The basics of machine learning
How to perform cross-validation to avoid over-training and how to build a recommendation system
What is regularization and why it is useful
Your Instructor
Professor of Biostatistics, Harvard T.H. Chan School of Public Health
Rafael Irizarry is a Professor of Biostatistics at the Harvard T.H. Chan School of Public Health and a Professor of Biostatistics and Computational Biology at the Dana Farber Cancer Institute. For the past 15 years, Dr. Irizarry’s research has focused on the analysis of genomics data. During this time, he has also taught several classes, all related to applied statistics. Dr. Irizarry is one of the founders of the Bioconductor Project, an open source and open development software project for the analysis of genomic data. His publications related to these topics have been highly cited and his software implementations widely downloaded.
Read full bio.
These courses can be bundled together to receive a professional certificate at a discounted price.
Learn More15-20 hours a week • Start today
1-2 hours a week • Start today
1-2 hours per week • Start today
1-2 hours a week • Start today
1-2 hours a week • Start today
1-2 hours a week • Start today
1-2 hours a week • Start today
1-2 hours a week • Start today
Ways to take this course
A Verified Certificate costs $149 and provides unlimited access to full course materials, activities, tests, and forums. At the end of the course, learners who earn a passing grade can receive a certificate.
Alternatively, learners can Audit the course for free and have access to select course material, activities, tests, and forums. Please note that this track does not offer a certificate for learners who earn a passing grade.
Don’t miss a thing. Subscribe to our newsletter and get updates on exclusive content for Harvard Online learners.