Data Science & Machine Learning Fundamentals
Machine learning is easy to buy and hard to judge. A vendor demonstration shows a model with 94 percent accuracy and the room nods, when the honest questions are what the base rate was, which errors the model makes, and whether the data it was trained on resembles the data it will meet. This day builds the understanding needed to ask those questions, and to work out which of your own problems are genuinely machine learning problems rather than reporting problems in disguise.
Programme Agenda
What Machine Learning Is Doing
Learning a pattern from examples instead of following written rules. Where that is the right approach and where a rule, a lookup table or a better report would do the job faster and more reliably.
Supervised, Unsupervised and the Rest
Classification and regression, clustering and anomaly detection, and where recommendation and forecasting sit. Matching a business question to a problem type.
The Data Requirement
How much data, of what quality, with what labels. Class imbalance, leakage from a column that would not exist at prediction time, and the historical bias a model will happily learn and reproduce.
Training, Validation and Testing
Why a model must be judged on data it has not seen, train and test splits, cross validation, and the time-based split that time series problems require.
Overfitting and Underfitting
The central failure mode, shown rather than described. Why a model that scores perfectly in development can be worthless in production, and the signals that give it away early.
Metrics That Match the Decision
Why accuracy misleads on imbalanced problems. Precision, recall, the trade between them, and choosing a threshold based on what a false positive and a false negative actually cost your business.
Interpretability and Trust
Feature importance, simple models against complex ones, and the regulatory and practical situations where you must be able to explain a decision to the person it affected.
Use Case Selection
Participants assess their own candidate use cases against data availability, decision value and error tolerance, and separate the promising ones from the ones that need a report instead.
Learning Outcomes:
Distinguish problems suited to machine learning from problems suited to rules
Match a business question to a problem type
Assess whether you have the data quantity, quality and labels required
Explain why models are evaluated on unseen data
Recognise overfitting and the signals that reveal it before deployment
Choose evaluation metrics based on the cost of each error type
Judge when interpretability is a requirement rather than a preference
Screen your own use cases for feasibility and value
Duration: 1 Day (8 Hours)
Training Hours: 9:00 AM to 5:00 PM
Level: Beginner
Training Mode: Physical, Online, or Hybrid
HRD Corp SBL-KHAS Claimable
Certificate of Completion included
Frequently Asked Questions
More in Data Engineering and Machine Learning
- MLOps and Model Deployment Basics · Beginner, 1 day
- Cloud Platform Foundations (AWS and Azure) · Beginner, 1 day
- Data Engineering and Pipelines · All levels, 2 days