DATA SCIENTIST COURSE

DATA SCIENTIST COURSE

Data Scientist Course

From Scratch to Essentials

Course Objectives

By the end of this course, students will be able to:

  • Understand the role of a Data Scientist and how it differs from a Data Analyst or Data Engineer
  • Work confidently with Python, NumPy, and Pandas for data manipulation
  • Clean, explore, and visualise data using EDA techniques
  • Understand core statistics and probability concepts used in data science
  • Perform basic hypothesis testing and interpret p-values and confidence intervals
  • Understand the machine learning workflow, including train/test splits
  • Build and evaluate simple regression and classification models
  • Understand basic unsupervised learning and clustering
  • Evaluate model performance using appropriate metrics
  • Complete a small end-to-end data science project

Course Syllabus

Day 1 — Introduction to Data Science. Covers what data science is, how the Data Scientist role differs from a Data Analyst or Data Engineer, and the end-to-end data science workflow (ask a question → collect data → clean → explore → model → communicate). Students are introduced to the tools used across the course (Python, NumPy, Pandas, Scikit-learn) and explore a sample dataset to see how a data science question moves from raw data to an actionable answer.

Day 2 — Python for Data Science. Covers Python fundamentals needed for data science — variables, data types, conditionals, loops, functions, and working with lists and dictionaries. Students write small scripts to process a sample dataset manually, building the foundation before moving into specialised libraries.

Day 3 — NumPy Fundamentals. Covers NumPy arrays, array creation, indexing and slicing, mathematical operations, broadcasting, and basic statistical functions (mean, median, standard deviation). Students practice performing vectorised calculations on numeric datasets, comparing performance and syntax against plain Python loops.

Day 4 — Pandas for Data Science. Covers DataFrames and Series, reading data from CSV/Excel, selecting and filtering rows/columns, handling missing values, and basic transformations (renaming, changing types, creating calculated columns). Students load and manipulate a real-world sample dataset using Pandas.

Day 5 — Data Cleaning & Exploratory Data Analysis (EDA). Covers identifying and handling missing values, duplicates, and outliers, along with generating summary statistics (describe(), value counts) to understand a dataset’s shape and quality. Students perform a full EDA pass on a messy sample dataset and document their initial findings.

Day 6 — Data Visualisation. Covers visualisation principles and building charts with Matplotlib and Seaborn — histograms, box plots, scatter plots, bar charts, and correlation heatmaps. Students visualise distributions and relationships in a sample dataset to identify patterns before modelling.

Day 7 — Statistics Fundamentals. Covers descriptive statistics — mean, median, mode, variance, standard deviation, percentiles, and the difference between population and sample statistics. Students calculate and interpret these measures on a real dataset to understand what the numbers reveal about the underlying data.

Day 8 — Probability Basics. Covers foundational probability concepts, independent vs. dependent events, and common distributions (normal, binomial) with an introduction to the Central Limit Theorem. Students explore how a normal distribution appears in real data and why it matters for later statistical testing.

Day 9 — Inferential Statistics & Hypothesis Testing. Covers confidence intervals, null and alternative hypotheses, p-values, and basic hypothesis tests (t-test). Students perform a simple hypothesis test on sample data — for example, testing whether two groups have significantly different average values — and interpret the results in plain language.

Day 10 — Introduction to Machine Learning. Covers what machine learning is, supervised vs. unsupervised learning, common use cases, and the standard ML workflow (train/test split, fitting a model, making predictions, evaluating results). Students split a sample dataset into training and testing sets in preparation for building their first models.

Day 11 — Regression. Covers simple and multiple linear regression, interpreting coefficients, and evaluating fit using R² and error metrics (MAE, RMSE). Students build a linear regression model to predict a numeric outcome (e.g., predicting sales from advertising spend) using Scikit-learn.

Day 12 — Classification. Covers classification concepts and building models with logistic regression, K-Nearest Neighbors (KNN), and an introduction to decision trees. Students train a classification model to predict a categorical outcome (e.g., whether a customer will churn) and generate predictions on test data.

Day 13 — Model Evaluation. Covers evaluation metrics for classification (accuracy, precision, recall, F1-score, confusion matrix) and regression (MAE, RMSE, R²), along with an introduction to cross-validation and the concepts of overfitting and underfitting. Students evaluate the models built on Days 11–12 and compare performance using these metrics.

Day 14 — Unsupervised Learning & Clustering. Covers unsupervised learning concepts and introduces K-Means clustering — how it works, choosing the number of clusters (elbow method), and interpreting cluster results. Students apply K-Means to a sample dataset (e.g., grouping customers by purchasing behaviour) and visualise the resulting clusters.

Day 15 — Final Data Science Project. Students receive a raw business dataset and complete an end-to-end mini project: (1) Clean & Explore — handle missing values/outliers and perform EDA with visualisations; (2) Statistical Analysis — calculate descriptive statistics and run a basic hypothesis test; (3) Modelling — build and train an appropriate regression or classification model; (4) Evaluation — assess model performance using the correct metrics; (5) Insights — identify key findings and business recommendations. The course closes with each student presenting their notebook and explaining: “What question did you ask, what does the data/model show, and what should the business do next?”


Core Topics Reference

Python & Libraries

  • Python basics: variables, conditionals, loops, functions
  • NumPy: arrays, indexing, broadcasting, statistical functions
  • Pandas: DataFrames, cleaning, filtering, transformations
  • Matplotlib / Seaborn: histograms, box plots, scatter plots, heatmaps

Statistics & Probability

  • Descriptive statistics: mean, median, mode, variance, standard deviation
  • Probability basics, normal & binomial distributions, Central Limit Theorem
  • Inferential statistics: confidence intervals, hypothesis testing, p-values, t-tests

Machine Learning

  • ML workflow: train/test split, fitting, predicting, evaluating
  • Regression: linear regression, R², MAE, RMSE
  • Classification: logistic regression, KNN, decision trees
  • Model evaluation: accuracy, precision, recall, F1-score, confusion matrix, cross-validation
  • Unsupervised learning: K-Means clustering, elbow method

Assessment Structure

Component Weight
Python / NumPy / Pandas Exercises 20%
Statistics & EDA Exercises 20%
Machine Learning Model Exercises 20%
Final Project 30%
Presentation 10%
Total 100%

Final Course Outcome

By the end of the 15-day course, students should have a practical, working data science flow:

Raw Data → Cleaning & EDA → Statistical Analysis → Model Building → Model Evaluation → Insights & Recommendations

Students should be able to take a raw dataset, explore and clean it, apply basic statistical reasoning, build a simple predictive model, evaluate it properly, and explain the results in business terms.

Recommended Session Structure

For each 2-hour session:

  • 20–30 minutes — Concept explanation
  • 60–70 minutes — Guided practical work
  • 20–30 minutes — Student exercise / assignment

This keeps the course hands-on and practical rather than theory-heavy, so students build real intuition for statistics and modelling rather than memorising formulas.

Total Fees: 90,000/=

Total Duration: 30 Hrs (2 hours  x  15 Classes)
Training Mode: Individual Training, your own timetable

Live Online Classes  or Face to Face Direct Classes with our expert trainers.
Call +94 (0) 722000999 / +94 (0) 755123111 www.iss.lk. Medium : සිංහල / தமிழ் / English