The Ultimate Hands-On Hadoop

Diese kurs ist nicht verfügbar in Deutsch (Deutschland)

Wir übersetzen es in weitere Sprachen.

The Ultimate Hands-On Hadoop

Dozent: Packt - Course Instructors

Bei Coursera Plus enthalten

Mehr erfahren

12 Module

Verschaffen Sie sich einen Einblick in ein Thema und lernen Sie die Grundlagen.

Stufe Mittel

Empfohlene Erfahrung

Es dauert 16 Stunden

3 Wochen bei 5 Stunden pro Woche

Flexibler Zeitplan

In Ihrem eigenen Lerntempo lernen

12 Module

Verschaffen Sie sich einen Einblick in ein Thema und lernen Sie die Grundlagen.

Stufe Mittel

Empfohlene Erfahrung

Es dauert 16 Stunden

3 Wochen bei 5 Stunden pro Woche

Flexibler Zeitplan

In Ihrem eigenen Lerntempo lernen

Was Sie lernen werden

Remember Hadoop setup and configuration steps.
Understand the Hadoop ecosystem, including HDFS, MapReduce, and YARN.
Apply queries using Pig, Hive, and Spark.
Evaluate Hadoop cluster performance and optimize it.

Kompetenzen, die Sie erwerben

Kategorie: MongoDB
Kategorie: spark
Kategorie: Hadoop
Kategorie: Kafka
Kategorie: Apache Hadoop
Kategorie: Big Data
Kategorie: Spark

Wichtige Details

Zertifikat zur Vorlage

Zu Ihrem LinkedIn-Profil hinzufügen

Kürzlich aktualisiert!

Oktober 2024

Bewertungen

5 Aufgaben

Unterrichtet in Englisch

Erfahren Sie, wie Mitarbeiter führender Unternehmen gefragte Kompetenzen erwerben.

Weitere Informationen zu Coursera für Unternehmen

Erwerben Sie ein Karrierezertifikat.

Fügen Sie diese Qualifikation zur Ihrem LinkedIn-Profil oder Ihrem Lebenslauf hinzu.

Teilen Sie es in den sozialen Medien und in Ihrer Leistungsbeurteilung.

In diesem Kurs gibt es 12 Module

Immerse yourself in the comprehensive world of Hadoop with this expertly designed course. Starting with the basics, you'll learn to install the Hortonworks Data Platform Sandbox on your local machine, providing you with a powerful environment to explore Hadoop's core functionalities. The course meticulously guides you through essential concepts such as the Hadoop Distributed File System (HDFS) and MapReduce, offering practical exercises to solidify your understanding.

As you progress, you'll delve into advanced Hadoop programming with tools like Pig, Hive, and Spark. These modules are designed to give you hands-on experience with real-world datasets, allowing you to build complex queries, analyze large datasets, and even venture into machine learning with Spark's MLLib. The course also covers integrating relational and non-relational databases with Hadoop, ensuring you can handle a wide range of data scenarios in your career. The final sections focus on managing and optimizing your Hadoop cluster, introducing you to tools like YARN, ZooKeeper, Oozie, and Kafka. You’ll learn how to feed data into your cluster efficiently, manage resources, and analyze streaming data in real time. By the end of this course, you’ll be well-equipped to design and implement Hadoop-based solutions in any data-driven environment. This course is ideal for data engineers, software developers, and IT professionals who have a basic understanding of programming and data management. Familiarity with Java, SQL, and Linux command-line interfaces is recommended but not required.

In this module, we will dive into the world of Hadoop, starting with its installation and setup using the Hortonworks Data Platform Sandbox. You'll explore the key buzzwords and technologies that make up the Hadoop ecosystem, learn about the historical context and impact of the Hortonworks and Cloudera merger, and begin working with real data to get a feel for Hadoop's capabilities.

Das ist alles enthalten

4 Videos1 Lektüre

In this module, we will explore the core components of Hadoop: the Hadoop Distributed File System (HDFS) and MapReduce. You'll learn how HDFS reliably stores massive data sets across a cluster and how MapReduce enables distributed data processing. Through hands-on activities, you'll import datasets, set up a MapReduce environment, and write scripts to analyze data, including breaking down movie ratings and ranking movies by popularity.

Das ist alles enthalten

10 Videos

10 VideosInsgesamt 94 Minuten

Hadoop Distributed File System (HDFS): What it is and How it Works13 MinutenModulvorschau
Installing the MovieLens Dataset6 Minuten
Activity - Installing the MovieLens Dataset into Hadoop's Distributed File System (HDFS) using the Command Line7 Minuten
MapReduce: What it is and How it Works10 Minuten
How MapReduce Distributes Processing12 Minuten
MapReduce Example: Breaking Down the Movie Ratings by Rating Score11 Minuten
Activity - Installing Python, MRJob, and Nano7 Minuten
Activity - Coding Up and Running the Ratings Histogram MapReduce Job7 Minuten
Exercise - Ranking Movies by Their Popularity7 Minuten
Activity - Checking Results8 Minuten

In this module, we will delve into Pig, a high-level scripting language that simplifies Hadoop programming. You'll start by exploring the Ambari web-based UI, which makes working with Pig more accessible. The module includes practical examples and activities, such as finding the oldest five-star movies and identifying the most-rated one-star movies using Pig scripts. You'll also learn about the capabilities of Pig Latin and test your skills through challenges and result comparisons.

Das ist alles enthalten

7 Videos1 Aufgabe

7 VideosInsgesamt 56 Minuten

Introducing Ambari9 MinutenModulvorschau
Introducing the Pig6 Minuten
Example - Finding the Oldest Movie with Five-Star Rating Using the Pig15 Minuten
Activity - Finding the Old Five-Star Movies with Pig9 Minuten
More Pig Latin7 Minuten
Exercise - Finding the Most-Rated One-Star Movie1 Minute
Pig Challenge - Comparing Results5 Minuten

1 AufgabeInsgesamt 15 Minuten

Assessment 115 Minuten

In this module, we will explore the power of Apache Spark, a key technology in the Hadoop ecosystem known for its speed and versatility. You’ll start by understanding why Spark is a game-changer in big data. The module will cover Resilient Distributed Datasets (RDDs) and Datasets, showing you how to use them to analyze movie ratings data. You'll also delve into Spark's machine learning library (MLLib) to create a movie recommendation system. Through hands-on activities, you'll practice writing Spark scripts and refining your data analysis skills.

Das ist alles enthalten

8 Videos

8 VideosInsgesamt 74 Minuten

Why Spark?10 MinutenModulvorschau
The Resilient Distributed Datasets (RDD)10 Minuten
Activity - Finding the Movie with the Lowest Average Rating with the Resilient Distributed Datasets (RDD)15 Minuten
Datasets and Spark 2.06 Minuten
Activity - Finding the movie with the Lowest Average Rating with DataFrames10 Minuten
Activity - Recommending a Movie with Spark's Machine Learning Library (MLLib)12 Minuten
Exercise - Filtering the Lowest-Rated Movies by Number of Ratings2 Minuten
Activity - Checking Results6 Minuten

In this module, we will explore the integration of relational datastores with Hadoop, focusing on Apache Hive and MySQL. You'll start by learning how Hive enables SQL queries on data within HDFS, followed by hands-on activities to find popular and highly-rated movies using Hive. The module also covers the installation and integration of MySQL with Hadoop, using Sqoop to seamlessly transfer data between MySQL and Hadoop's HDFS/Hive. Through practical exercises, you'll gain proficiency in managing and querying relational data within the Hadoop ecosystem.

Das ist alles enthalten

9 Videos

9 VideosInsgesamt 63 Minuten

What is Hive?6 MinutenModulvorschau
Activity - Using Hive to Find the Most Popular Movie10 Minuten
How Hive Works?9 Minuten
Exercise - Using Hive to Find the Movie with the Highest Average Rating1 Minute
Comparing Solutions4 Minuten
Integrating MySQL with Hadoop8 Minuten
Activity - Installing MySQL and Importing Movie Data7 Minuten
Activity - Using Sqoop to Import Data from MySQL to HFDS/Hive7 Minuten
Activity - Using Sqoop to Export Data from Hadoop to MySQL7 Minuten

In this module, we will explore the use of non-relational (NoSQL) data stores within the Hadoop ecosystem. You'll learn why NoSQL databases are crucial for scalability and efficiency, and dive into specific technologies like HBase, Cassandra, and MongoDB. Through a series of activities, you'll practice importing data into HBase, integrating it with Pig, and using Cassandra and MongoDB alongside Spark. The module concludes with exercises to help you choose the most suitable NoSQL database for different scenarios, empowering you to make informed decisions in big data management.

Das ist alles enthalten

12 Videos1 Aufgabe

12 VideosInsgesamt 147 Minuten

Why NoSQL?13 MinutenModulvorschau
What is HBase?12 Minuten
Activity - Importing Movie Ratings into HBase13 Minuten
Activity - Using HBase with Pig to Import Data at Scale11 Minuten
Cassandra - Overview14 Minuten
Activity - Installing Cassandra11 Minuten
Activity - Writing Spark Output into Cassandra11 Minuten
MongoDB - Overview17 Minuten
Activity - Installing MongoDB and Integrating Spark with MongoDB12 Minuten
Activity - Using the MongoDB Shell7 Minuten
Choosing Database Technology15 Minuten
Exercise - Choosing a Database for a Given Problem5 Minuten

1 AufgabeInsgesamt 15 Minuten

Assessment 215 Minuten

In this module, we will focus on interactive querying tools that allow you to quickly access and analyze big data across multiple sources. You'll explore technologies like Drill, Phoenix, and Presto, learning how each one solves specific challenges in querying large datasets. The module includes hands-on activities where you'll set up these tools, execute queries that span across databases such as MongoDB, Hive, HBase, and Cassandra, and integrate these tools with other Hadoop ecosystem components. By the end of this module, you'll be equipped to perform efficient, real-time data analysis across varied data stores.

Das ist alles enthalten

9 Videos

9 VideosInsgesamt 81 Minuten

Overview of Drill7 MinutenModulvorschau
Activity - Setting Up Drill10 Minuten
Activity - Querying Across Multiple Databases with Drill7 Minuten
Overview of Phoenix8 Minuten
Activity - Installing Phoenix and Querying HBase7 Minuten
Activity - Integrating Phoenix with the Pig11 Minuten
Overview of Presto6 Minuten
Activity - Installing Presto and Querying Hive12 Minuten
Activity - Querying Both Cassandra and Hive Using Presto9 Minuten

In this module, we will explore the critical components involved in managing a Hadoop cluster. You'll learn about YARN's resource management capabilities, how Tez optimizes task execution using Directed Acyclic Graphs, and the differences between Mesos and YARN. We'll dive into ZooKeeper for maintaining reliable operations and Oozie for orchestrating complex workflows. Hands-on activities will guide you through setting up and using Zeppelin for interactive data analysis and using Hue for a more user-friendly interface. The module also touches on other noteworthy technologies like Chukwa and Ganglia, providing a comprehensive understanding of cluster management in Hadoop.

Das ist alles enthalten

13 Videos

13 VideosInsgesamt 119 Minuten

Yet Another Resource Negotiator (YARN)10 MinutenModulvorschau
Tez4 Minuten
Activity - Using Hive on Tez and Measuring the Performance Benefit8 Minuten
Mesos7 Minuten
ZooKeeper13 Minuten
Activity - Simulating a Failing Master with ZooKeeper6 Minuten
Oozie11 Minuten
Activity - Setting Up a Simple Oozie Workflow16 Minuten
Zeppelin - Overview5 Minuten
Hands-On with Zeppelin for Spark and MovieLens Analysis12 Minuten
SQL and Data Visualization in Zeppelin: MovieLens Analysis with Spark9 Minuten
Hue - Overview8 Minuten
Other Technologies Worth Mentioning4 Minuten

In this module, we will explore the essential tools for feeding data into your Hadoop cluster, focusing on Kafka and Flume. You'll learn how Kafka supports scalable and reliable data collection across a cluster and how to set it up to publish and consume data. Additionally, you'll discover how Flume's architecture differs from Kafka and how to use it for real-time data ingestion. Through hands-on activities, you'll configure Kafka to monitor Apache logs and Flume to watch directories, publishing incoming data into HDFS. These skills will help you manage and process streaming data effectively in your Hadoop environment.

Das ist alles enthalten

6 Videos1 Aufgabe

6 VideosInsgesamt 54 Minuten

Kafka9 MinutenModulvorschau
Activity - Setting Up Kafka and Publishing Data7 Minuten
Activity - Publishing Web Logs with Kafka10 Minuten
Flume10 Minuten
Activity - Setting up Flume and Publishing Logs7 Minuten
Activity - Setting Up Flume to Monitor a Directory and Store its Data in Hadoop Distributed File System (HDFS)9 Minuten

1 AufgabeInsgesamt 15 Minuten

Assessment 315 Minuten

In this module, we will focus on analyzing streams of data using real-time processing frameworks such as Spark Streaming, Apache Storm, and Flink. You’ll start by learning how Spark Streaming processes micro-batches of data in real-time and participate in activities that include analyzing web logs streamed by Flume. The module then introduces Apache Storm and Flink, providing hands-on exercises to implement word count applications with these tools. By the end of this module, you will be able to build continuous applications that efficiently process and analyze streaming data.

Das ist alles enthalten

8 Videos

8 VideosInsgesamt 76 Minuten

Spark Streaming: Introduction14 MinutenModulvorschau
Activity - Analyzing Web Logs Published with Flume using Spark Streaming14 Minuten
Exercise - Monitor Flume-Published Logs for Errors in Real Time2 Minuten
Exercise Solution: Aggregating the Hypertext Transfer Protocol (HTTP) Access Codes with Spark Streaming4 Minuten
Apache Storm: Introduction9 Minuten
Activity - Counting Words with Storm14 Minuten
Flink: Overview6 Minuten
Activity - Counting Words with Flink10 Minuten

In this module, we will focus on designing and implementing real-world systems using a combination of Hadoop ecosystem tools. You'll start by exploring additional technologies like Impala, NiFi, and AWS Kinesis, learning how they fit into broader Hadoop-based solutions. The module then guides you through the process of understanding system requirements and designing applications that consume and analyze large-scale data, such as web server logs or movie recommendations. By the end of this module, you’ll be equipped to design and build complex, efficient, and scalable data systems tailored to specific business needs.

Das ist alles enthalten

7 Videos1 Aufgabe

7 VideosInsgesamt 52 Minuten

The Best of the Rest9 MinutenModulvorschau
Review: How the Pieces Fit Together?6 Minuten
Understanding Your Requirements8 Minuten
Sample Application: Consuming Web Server Logs and Keeping Track of Top-Sellers10 Minuten
Sample Application: Serving Movie Recommendations to a Website11 Minuten
Exercise - Designing a System to Report Web Sessions Per Day2 Minuten
Exercise Solution: Designing a System to Count Daily Sessions4 Minuten

1 AufgabeInsgesamt 15 Minuten

Assessment 415 Minuten

In this final module, we will provide you with a selection of books, online resources, and tools recommended by the author to further your knowledge of Hadoop and related technologies. This module serves as a guide for continued learning, offering you the means to stay updated with the latest developments in the Hadoop ecosystem and expand your skills beyond this course.

Das ist alles enthalten

1 Video1 Aufgabe

Dozent

Packt - Course Instructors

Packt

375 Kurse25.243 Lernende

von

Packt

Warum entscheiden sich Menschen für Coursera für ihre Karriere?

Felipe M.

Lernender seit 2018

„Es ist eine großartige Erfahrung, in meinem eigenen Tempo zu lernen. Ich kann lernen, wenn ich Zeit und Nerven dazu habe.“

Jennifer J.

Lernender seit 2020

„Bei einem spannenden neuen Projekt konnte ich die neuen Kenntnisse und Kompetenzen aus den Kursen direkt bei der Arbeit anwenden.“

Larry W.

Lernender seit 2021

„Wenn mir Kurse zu Themen fehlen, die meine Universität nicht anbietet, ist Coursera mit die beste Alternative.“

Chaitanya A.

„Man lernt nicht nur, um bei der Arbeit besser zu werden. Es geht noch um viel mehr. Bei Coursera kann ich ohne Grenzen lernen.“

Neue Karrieremöglichkeiten mit Coursera Plus

Unbegrenzter Zugang zu 10,000+ Weltklasse-Kursen, praktischen Projekten und berufsqualifizierenden Zertifikatsprogrammen - alles in Ihrem Abonnement enthalten

Mehr erfahren

Bringen Sie Ihre Karriere mit einem Online-Abschluss voran.

Erwerben Sie einen Abschluss von erstklassigen Universitäten – 100 % online

Erkunden Sie die Abschlüsse

Schließen Sie sich mehr als 3.400 Unternehmen in aller Welt an, die sich für Coursera for Business entschieden haben.

Schulen Sie Ihre Mitarbeiter*innen, um sich in der digitalen Wirtschaft zu behaupten.

Mehr erfahren

Häufig gestellte Fragen

Yes, you can preview the first video and view the syllabus before you enroll. You must purchase the course to access content not included in the preview.

If you decide to enroll in the course before the session start date, you will have access to all of the lecture videos and readings for the course. You’ll be able to submit assignments once the session starts.

Once you enroll and your session begins, you will have access to all videos and other resources, including reading items and the course discussion forum. You’ll be able to view and submit practice assessments, and complete required graded assignments to earn a grade and a Course Certificate.

Weitere Fragen

Besuchen Sie die das Hilfe-Center für Kursteilnehmer.

Diese kurs ist nicht verfügbar in Deutsch (Deutschland)

The Ultimate Hands-On Hadoop

Was Sie lernen werden

Kompetenzen, die Sie erwerben

Wichtige Details

Erfahren Sie, wie Mitarbeiter führender Unternehmen gefragte Kompetenzen erwerben.

Erwerben Sie ein Karrierezertifikat.

In diesem Kurs gibt es 12 Module

Learning All the Buzzwords and Installing the Hortonworks Data Platform Sandbox

Das ist alles enthalten

Using the Hadoop's Core: Hadoop Distributed File System (HDFS) and MapReduce

Das ist alles enthalten

Programming Hadoop with Pig

Das ist alles enthalten

Programming Hadoop with Spark

Das ist alles enthalten

Using Relational Datastores with Hadoop

Das ist alles enthalten

Using Non-Relational Data Stores with Hadoop

Das ist alles enthalten

Querying Data Interactively

Das ist alles enthalten

Managing Your Cluster

Das ist alles enthalten

Feeding Data to Your Cluster

Das ist alles enthalten

Analyzing Streams of Data

Das ist alles enthalten

Designing Real-World Systems

Das ist alles enthalten

Learning More

Das ist alles enthalten

Dozent

von

Empfohlen, wenn Sie sich für Data Management interessieren

Kafka Integration with Storm, Spark, Flume, and Security

SQL: A Practical Introduction for Querying Databases

Data Engineering Capstone Project

Big Data: capstone project

Warum entscheiden sich Menschen für Coursera für ihre Karriere?

Neue Karrieremöglichkeiten mit Coursera Plus

Bringen Sie Ihre Karriere mit einem Online-Abschluss voran.

Schließen Sie sich mehr als 3.400 Unternehmen in aller Welt an, die sich für Coursera for Business entschieden haben.

Häufig gestellte Fragen

Can I preview a course before enrolling?

When will I have access to the lectures and assignments?

What will I get when I enroll?

Weitere Fragen