Course Outline
Introduction:
- The Role of Apache Spark in the Hadoop Ecosystem
- Brief Overview of Python and Scala
Core Concepts (Theory):
- System Architecture
- Resilient Distributed Datasets (RDD)
- Transformations and Actions
- Stages, Tasks, and Dependencies
Practical Workshop: Mastering Basics in the Databricks Environment:
- Hands-on exercises with the RDD API
- Fundamental action and transformation functions
- PairRDDs
- Join operations
- Caching strategies
- Hands-on exercises with the DataFrame API
- SparkSQL
- DataFrame operations: select, filter, groupBy, and sortBy
- User-Defined Functions (UDFs)
- Exploring the DataSet API
- Streaming capabilities
Practical Workshop: Deployment Strategies in the AWS Environment:
- Fundamentals of AWS Glue
- Differences between AWS EMR and AWS Glue
- Example jobs executed on both platforms
- Advantages and disadvantages of each approach
Additional Content:
- Introduction to Apache Airflow for orchestration
Requirements
Programming skills (preferably in Python or Scala)
Basic knowledge of SQL
Custom Corporate Training
Training solutions designed exclusively for businesses.
- Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
- Flexible Schedule: Dates and times adapted to your team's agenda.
- Format: Online (live), In-company (at your offices), or Hybrid.
Price per private group, online live training, starting from 4800 € + VAT*
Contact us for an exact quote and to hear our latest promotions
Testimonials (3)
Having hands on session / assignments
Poornima Chenthamarakshan - Intelligent Medical Objects
Course - Apache Spark in the Cloud
1. Right balance between high level concepts and technical details. 2. Andras is very knowledgeable about his teaching. 3. Exercise
Steven Wu - Intelligent Medical Objects
Course - Apache Spark in the Cloud
Get to learn spark streaming , databricks and aws redshift