Mathew K Analytics

Lesson 3 · Python For Machine Learning

Comprehensive Guide to Training Machine Learning Models with Scikit-Learn in Python

Welcome! In this lesson, we will explore what machine learning is, how it uses data, and how to get started with simple Python tools. By the end, you will…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Introduction to Machine Learning in Python#

Welcome! In this lesson, we will explore what machine learning is, how it uses data, and how to get started with simple Python tools.

By the end, you will know how to work with a real-world dataset and run a simple machine learning task.

# Import the warnings library to keep our notebook clean
import warnings
warnings.filterwarnings('ignore')

# We are setting this up to get fewer distracting messages.

What is Machine Learning?#

Machine learning means teaching computers to learn from data like a calculator that can spot patterns and make predictions.

We use simple code to train models and let them solve problems for us.

# Machine learning always starts with data.
# Let's use the famous Titanic dataset.

import pandas as pd
titanic = pd.read_csv("https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv")

# Show how many rows and columns we have.
print("Shape:", titanic.shape)

# See the top few rows.
titanic.head()
Shape: (891, 12)
PassengerId Survived Pclass Name Sex Age SibSp Parch Ticket Fare Cabin Embarked
0 1 0 3 Braund, Mr. Owen Harris male 22.0 1 0 A/5 21171 7.2500 NaN S
1 2 1 1 Cumings, Mrs. John Bradley (Florence Briggs Th... female 38.0 1 0 PC 17599 71.2833 C85 C
2 3 1 3 Heikkinen, Miss. Laina female 26.0 0 0 STON/O2. 3101282 7.9250 NaN S
3 4 1 1 Futrelle, Mrs. Jacques Heath (Lily May Peel) female 35.0 1 0 113803 53.1000 C123 S
4 5 0 3 Allen, Mr. William Henry male 35.0 0 0 373450 8.0500 NaN S
# Let us look at the columns in our data.
print(titanic.columns.tolist())

# Each column is a piece of information we can use.
['PassengerId', 'Survived', 'Pclass', 'Name', 'Sex', 'Age', 'SibSp', 'Parch', 'Ticket', 'Fare', 'Cabin', 'Embarked']
# Let us see how many people survived.
print(titanic['Survived'].value_counts())

# Survived is our target  this is what we want to predict.
Survived
0    549
1    342
Name: count, dtype: int64

How does a machine learn?#

First, we show the machine examples: rows from our dataset.

Then, we ask it to find patterns so it can make good guesses about new data.

Our example: Can the computer guess if a passenger survived, based on their information?

# Let us prepare our data for the computer.
# We will only use columns that have no missing values for now.
features = ['Pclass', 'Sex', 'Age', 'SibSp', 'Parch', 'Fare']

# Drop rows with missing values in these columns.
titanic_clean = titanic.dropna(subset=features)

# Fill in the Sex column as numbers: male is 0, female is 1.
titanic_clean['Sex'] = titanic_clean['Sex'].map({'male': 0, 'female': 1})
# Set up our feature data (X) and our target (y).
X = titanic_clean[features]
y = titanic_clean['Survived']
# Split our data into a train set and a test set.
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

print('Train shape:', X_train.shape, 'Test shape:', X_test.shape)
Train shape: (571, 6) Test shape: (143, 6)

What is a model?#

A model is a smart tool the computer builds to predict something.

It uses patterns it found in the data we gave it.

We will use a Decision Tree, a simple and visual kind of model.

# Make our model: a decision tree classifier.
from sklearn.tree import DecisionTreeClassifier
model = DecisionTreeClassifier(random_state=42)

# Fit (train) the model using training data.
model.fit(X_train, y_train)
DecisionTreeClassifier(random_state=42)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Use our trained model to make predictions.
y_pred = model.predict(X_test)

# Compare predictions to the true answers.

from sklearn.metrics import accuracy_score
acc = accuracy_score(y_test, y_pred)
print("Accuracy:", acc)
Accuracy: 0.7132867132867133
# Show predictions for the first 10 passengers in our test set.
print("Predicted:", y_pred[:10])
print("Actual:   ", y_test.values[:10])
Predicted: [1 1 1 1 0 0 0 0 0 1]
Actual:    [0 1 1 1 0 1 1 1 0 0]
# We can see which input features are most important to our model.
import matplotlib.pyplot as plt
importance = model.feature_importances_
plt.bar(features, importance)
plt.title("Feature Importance")
plt.show()
No description has been provided for this image
# What happens if you give the model unexpected input? Try running this cell:
try:
    model.predict([[3, 1, 30, 0, 0]])
except Exception as e:
    print("Error:", e)
    
Error: X has 5 features, but DecisionTreeClassifier is expecting 6 features as input.
# Let us let a user try their own passenger! Enter values when prompted.
pclass = int(input("Enter Pclass (1, 2, or 3): "))
sex = int(input("Enter Sex (0 for male, 1 for female): "))
age = float(input("Enter Age: "))
sibsp = int(input("Enter number of siblings/spouses aboard: "))
parch = int(input("Enter number of parents/children aboard: "))
fare = float(input("Enter Fare: "))

my_features = [[pclass, sex, age, sibsp, parch, fare]]
result = model.predict(my_features)
if result[0] == 1:
    print("The model predicts: Survived")
else:
    print("The model predicts: Did not survive")
    
The model predicts: Survived
 

Tips and Tricks#

  • If you get an error, read the message it often tells you what went wrong.
  • Try changing model settings, or using different features.
  • Remember: More data often helps the model improve.
# Try another model: logistic regression.
from sklearn.linear_model import LogisticRegression

log_model = LogisticRegression()
log_model.fit(X_train, y_train)

log_acc = log_model.score(X_test, y_test)
print("Logistic Regression accuracy:", log_acc)
Logistic Regression accuracy: 0.7482517482517482
# Challenge: What do you think will happen if you use only one feature?
X_age = X_train[['Age']]
model_one = DecisionTreeClassifier(random_state=42)
model_one.fit(X_age, y_train)
score_one = model_one.score(X_test[['Age']], y_test)
print('Accuracy with Age only:', score_one)
Accuracy with Age only: 0.6153846153846154
# Quick recap: What is in a machine learning pipeline?
# 1. Data loading
# 2. Data cleaning and prep
# 3. Split data train/test
# 4. Model building
# 5. Model testing
# 6. Use the model for new predictions
# Extra: Save your results to a CSV file.
output = X_test.copy()
output['Predicted'] = y_pred
output.to_csv('my_titanic_predictions.csv', index=False)
print('Saved to my_titanic_predictions.csv')
Saved to my_titanic_predictions.csv

Practice Time!#

Try changing which features you use, or swap in a different model.

Experiment and see how it changes predictions.

Challenge: Can you predict if a child with no family and a first-class ticket would survive?

If you enjoy this, subscribe and explore other machine learning datasets!

Lesson Recap#

  • You learned what machine learning is.
  • You loaded and explored real data.
  • You trained, tested, and used models!

You are ready to learn more!

Thank You!#

If you found this useful, please like and subscribe for more beginner Python and machine learning lessons.

See you in the next video!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.