This course will teach you how to leverage the power of Python and artificial intelligence to create and test hypothesis. We'll start for the ground up, learning some basic Python for data science before diving into some of its richer applications to test our created hypothesis. We'll learn some of the most important libraries for exploratory data analysis (EDA) and machine learning such as Numpy, Pandas, and Sci-kit learn. After learning some of the theory (and math) behind linear regression, we'll go through and full pipeline of reading data, cleaning it, and applying a regression model to estimate the progression of diabetes. By the end of the course, you'll apply a classification model to predict the presence/absence of heart disease from a patient's health data.
Introduction to Data Science and scikit-learn in Python
This course is part of AI for Scientific Research Specialization
Instructors: Sabrina Moore
Sponsored by Abu Dhabi National Oil Company
6,683 already enrolled
(45 reviews)
Recommended experience
What you'll learn
Employ artificial intelligence techniques to test hypothesis in Python
Apply a machine learning model combining Numpy, Pandas, and Scikit-Learn
Skills you'll gain
Details to know
Add to your LinkedIn profile
9 assignments
See how employees at top companies are mastering in-demand skills
Build your subject-matter expertise
- Learn new concepts from industry experts
- Gain a foundational understanding of a subject or tool
- Develop job-relevant skills with hands-on projects
- Earn a shareable career certificate
Earn a career certificate
Add this credential to your LinkedIn profile, resume, or CV
Share it on social media and in your performance review
There are 4 modules in this course
In this module, we'll get ourselves started with Programming in Python. After becoming familiar with Python and the Jupyter Notebook interface, we'll dive into some basic coding paradigms such as variables, loops, and functions. We'll also cover data structures in the form of lists and dictionaries. We'll go through one of the most useful things in your Python arsenal - importing and using modules effectively. Finally, we'll introduce scikit-learn and walk through a classification problem to predict the presence/absence of cancer from health data.
What's included
9 videos5 readings2 assignments4 programming assignments1 discussion prompt5 ungraded labs
In this module, we'll become familiar with the two most important packages for data science: Numpy and Pandas. We'll begin by learning the differences between the two packages. Then, we'll get ourselves familiar with np arrays and their functionalities. Adding text turns our arrays into tables, and gives rise to the Pandas module. After a basic introduction, we'll end with a series of important data manipulation tools such as indexing, merging/combining datasets, and reshaping data.
What's included
8 videos5 readings4 assignments1 programming assignment1 discussion prompt2 ungraded labs
In this module, we'll work from the ground up to build and test our hypothesis. Learning both the theory and the code, we'll learn to test our predictions with different types of machine learning algorithms. We'll start by going through some of the necessary data preprocessing steps to orient ourselves. Getting familiar with using the Scikit-Learn library starts with the documentation. From there, we'll load in a dataset and analyze some of its most basic properties. Finally, we'll import and use models to make a prediction.
What's included
6 videos3 readings3 assignments1 programming assignment1 discussion prompt1 ungraded lab
In the final project, we'll try and predict the presence of heart disease using patient data. We'll load in data, create new features, and apply a machine learning algorithm using scikit-learn.
What's included
1 video1 programming assignment1 ungraded lab
Offered by
Why people choose Coursera for their career
Learner reviews
45 reviews
- 5 stars
47.82%
- 4 stars
15.21%
- 3 stars
15.21%
- 2 stars
10.86%
- 1 star
10.86%
Showing 3 of 45
Reviewed on Nov 9, 2021
meskipun agak eror dalam lab penugasan tapi alhamdulillah sudah bisa
Reviewed on Nov 27, 2021
Good introduction. A bit too short for a 4-week course. The autograder is not very good, and some solutions are wrong.
Reviewed on Apr 4, 2022
The topic is great, and the linkage and references provided are valuable.
Recommended if you're interested in Data Science
University of Pennsylvania
Imperial College London
University of Colorado Boulder
Open new doors with Coursera Plus
Unlimited access to 10,000+ world-class courses, hands-on projects, and job-ready certificate programs - all included in your subscription
Advance your career with an online degree
Earn a degree from world-class universities - 100% online
Join over 3,400 global companies that choose Coursera for Business
Upskill your employees to excel in the digital economy