Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 21 hours (3 days)
Course Outline
Foundations of Audio Classification
- Categorization of sound events: environmental, mechanical, and human-generated.
- Overview of use cases: surveillance, monitoring, and automation.
- Distinguishing between audio classification, detection, and segmentation.
Audio Data and Feature Extraction
- Variations in audio file types and formats.
- Considerations for sampling rate, windowing, and frame size.
- Extraction of MFCCs, chroma features, and mel-spectrograms.
Data Preparation and Annotation
- Utilization of UrbanSound8K, ESC-50, and custom datasets.
- Labeling of sound events and their temporal boundaries.
- Strategies for balancing datasets and audio augmentation.
Developing Audio Classification Models
- Application of convolutional neural networks (CNNs) for audio analysis.
- Model input options: raw waveform versus extracted features.
- Selection of loss functions, evaluation metrics, and management of overfitting.
Event Detection and Temporal Localization
- Implementation of frame-based and segment-based detection strategies.
- Post-processing of detections utilizing thresholds and smoothing techniques.
- Visualization of predictions across audio timelines.
Advanced Topics and Real-Time Processing
- Application of transfer learning in scenarios with limited data.
- Deployment of models using TensorFlow Lite or ONNX.
- Streaming audio processing and evaluation of latency considerations.
Project Development and Application Scenarios
- Designing a comprehensive pipeline from data ingestion to classification.
- Creating a proof-of-concept for surveillance, quality control, or monitoring purposes.
- Integration of logging, alerting systems, and dashboards or APIs.
Summary and Next Steps
Requirements
- A solid grasp of machine learning concepts and the model training process.
- Practical experience with Python programming and data preprocessing workflows.
- Familiarity with the fundamentals of digital audio.
Intended Audience
- Data scientists.
- Machine learning engineers.
- Researchers and developers specializing in audio signal processing.
Custom Corporate Training
Training solutions designed exclusively for businesses.
- Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
- Flexible Schedule: Dates and times adapted to your team's agenda.
- Format: Online (live), In-company (at your offices), or Hybrid.
Price per private group, online live training, starting from 4800 € + VAT*
Contact us for an exact quote and to hear our latest promotions