Get in Touch
 Duration 21 hours (3 days)

Course Outline

Foundations of Audio Classification

  • Categorization of sound events: environmental, mechanical, and human-generated.
  • Overview of use cases: surveillance, monitoring, and automation.
  • Distinguishing between audio classification, detection, and segmentation.

Audio Data and Feature Extraction

  • Variations in audio file types and formats.
  • Considerations for sampling rate, windowing, and frame size.
  • Extraction of MFCCs, chroma features, and mel-spectrograms.

Data Preparation and Annotation

  • Utilization of UrbanSound8K, ESC-50, and custom datasets.
  • Labeling of sound events and their temporal boundaries.
  • Strategies for balancing datasets and audio augmentation.

Developing Audio Classification Models

  • Application of convolutional neural networks (CNNs) for audio analysis.
  • Model input options: raw waveform versus extracted features.
  • Selection of loss functions, evaluation metrics, and management of overfitting.

Event Detection and Temporal Localization

  • Implementation of frame-based and segment-based detection strategies.
  • Post-processing of detections utilizing thresholds and smoothing techniques.
  • Visualization of predictions across audio timelines.

Advanced Topics and Real-Time Processing

  • Application of transfer learning in scenarios with limited data.
  • Deployment of models using TensorFlow Lite or ONNX.
  • Streaming audio processing and evaluation of latency considerations.

Project Development and Application Scenarios

  • Designing a comprehensive pipeline from data ingestion to classification.
  • Creating a proof-of-concept for surveillance, quality control, or monitoring purposes.
  • Integration of logging, alerting systems, and dashboards or APIs.

Summary and Next Steps

Requirements

  • A solid grasp of machine learning concepts and the model training process.
  • Practical experience with Python programming and data preprocessing workflows.
  • Familiarity with the fundamentals of digital audio.

Intended Audience

  • Data scientists.
  • Machine learning engineers.
  • Researchers and developers specializing in audio signal processing.

Custom Corporate Training

Training solutions designed exclusively for businesses.

  • Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
  • Flexible Schedule: Dates and times adapted to your team's agenda.
  • Format: Online (live), In-company (at your offices), or Hybrid.
Investment

Price per private group, online live training, starting from 4800 € + VAT*

Contact us for an exact quote and to hear our latest promotions

Provisional Upcoming Courses (Contact Us For More Information)

Related Categories