Get in Touch
 Duration 21 hours

Course Outline

Comprehensive training syllabus

  1. Introduction to NLP
    • Core concepts of NLP
    • Overview of NLP frameworks
    • Commercial use cases for NLP
    • Web data scraping techniques
    • Retrieving text data via various APIs
    • Managing text corpora: storing content and associated metadata
    • Benefits of Python and an introductory NLTK session
  2. Practical Understanding of a Corpus and Dataset
    • The necessity of a corpus
    • Analyzing corpora
    • Varieties of data attributes
    • File formats for corpora
    • Preparing datasets for NLP workflows
  3. Understanding the Structure of Sentences
    • Fundamental NLP components
    • Natural language understanding
    • Morphological analysis: stems, words, tokens, and speech tags
    • Syntactic analysis
    • Semantic analysis
    • Managing ambiguity
  4. Text Data Preprocessing
    • Raw text corpus
      • Sentence tokenization
      • Stemming raw text
      • Lemmatizing raw text
      • Removing stop words
    • Raw sentence corpus
      • Word tokenization
      • Word lemmatization
    • Managing Term-Document and Document-Term matrices
    • Tokenizing text into n-grams and sentences
    • Customized and practical preprocessing strategies
  5. Analyzing Text Data
    • Foundational NLP features
      • Parsers and parsing techniques
      • POS tagging and taggers
      • Named entity recognition
      • N-grams
      • Bag of words
    • Statistical NLP features
      • Linear algebra concepts for NLP
      • Probabilistic theory in NLP
      • TF-IDF
      • Vectorization
      • Encoders and Decoders
      • Normalization
      • Probabilistic Models
    • Advanced feature engineering and NLP
      • word2vec fundamentals
      • Components of the word2vec model
      • The underlying logic of word2vec
      • Extending the word2vec concept
      • Applying the word2vec model
    • Case study: Applying bag of words to automatic text summarization using simplified and true Luhn's algorithms
  6. Document Clustering, Classification, and Topic Modeling
    • Document clustering and pattern mining (e.g., hierarchical clustering, k-means)
    • Document comparison and classification using TFIDF, Jaccard, and cosine distance
    • Document classification via Naïve Bayes and Maximum Entropy
  7. Identifying Important Text Elements
    • Dimensionality reduction: PCA, SVD, and Non-negative Matrix Factorization
    • Topic modeling and information retrieval using Latent Semantic Analysis
  8. Entity Extraction, Sentiment Analysis, and Advanced Topic Modeling
    • Distinguishing positive vs. negative sentiment degrees
    • Item Response Theory
    • Applying POS tagging to identify people, places, and organizations
    • Advanced topic modeling with Latent Dirichlet Allocation
  9. Case Studies
    • Mining unstructured user reviews
    • Sentiment classification and visualization of product review data
    • Analyzing search logs for usage patterns
    • Text classification
    • Topic modelling

Requirements

Familiarity with NLP principles and an understanding of how AI is applied in business contexts.

Custom Corporate Training

Training solutions designed exclusively for businesses.

  • Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
  • Flexible Schedule: Dates and times adapted to your team's agenda.
  • Format: Online (live), In-company (at your offices), or Hybrid.
Investment

Price per private group, online live training, starting from 4800 € + VAT*

Contact us for an exact quote and to hear our latest promotions

Testimonials (1)

Provisional Upcoming Courses (Contact Us For More Information)

Related Categories