Get in Touch

Course Outline

Introduction:

  • The Role of Apache Spark in the Hadoop Ecosystem
  • Brief Overview of Python and Scala

Core Concepts (Theory):

  • System Architecture
  • Resilient Distributed Datasets (RDD)
  • Transformations and Actions
  • Stages, Tasks, and Dependencies

Practical Workshop: Mastering Basics in the Databricks Environment:

  • Hands-on exercises with the RDD API
  • Fundamental action and transformation functions
  • PairRDDs
  • Join operations
  • Caching strategies
  • Hands-on exercises with the DataFrame API
  • SparkSQL
  • DataFrame operations: select, filter, groupBy, and sortBy
  • User-Defined Functions (UDFs)
  • Exploring the DataSet API
  • Streaming capabilities

Practical Workshop: Deployment Strategies in the AWS Environment:

  • Fundamentals of AWS Glue
  • Differences between AWS EMR and AWS Glue
  • Example jobs executed on both platforms
  • Advantages and disadvantages of each approach

Additional Content:

  • Introduction to Apache Airflow for orchestration

Requirements

Programming skills (preferably in Python or Scala)

Basic knowledge of SQL

 21 Hours

Custom Corporate Training

Training solutions designed exclusively for businesses.

  • Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
  • Flexible Schedule: Dates and times adapted to your team's agenda.
  • Format: Online (live), In-company (at your offices), or Hybrid.
Investment

Price per private group, online live training, starting from 4800 € + VAT*

Contact us for an exact quote and to hear our latest promotions

Testimonials (3)

Provisional Upcoming Courses (Contact Us For More Information)

Related Categories