Mathew K Analytics

Lesson 1 · Python For Machine Learning

1 Capstone ML Project in Python: Build and Evaluate a Machine Learning Model

Welcome! You are going to learn how to build a real machine learning project using Python. No coding background needed. We will guide you step by step with…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb
 

Capstone Project: Build Your First ML Model in Python#

Welcome! You are going to learn how to build a real machine learning project using Python.

No coding background needed. We will guide you step by step with small, friendly examples.

You will work with data, make predictions, and see how machine learning can solve real-world problems.

Ready? Let us dive in!

# Let us import some helpful libraries.
import warnings
warnings.filterwarnings("ignore")  # We will ignore warning messages to keep it clean.
import pandas as pd
import numpy as np

What is Machine Learning?#

Machine learning is a way for computers to learn from data, like people learn from experience.

You give the computer data and some examples of the answers it should learn.

Then, it can make predictions about new data.

Today, we will use the Titanic dataset to predict who survived the famous shipwreck.

# Data setup
url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv"
titanic = pd.read_csv(url)
print("Rows and columns:", titanic.shape)
titanic.head()
Rows and columns: (891, 12)
PassengerId Survived Pclass Name Sex Age SibSp Parch Ticket Fare Cabin Embarked
0 1 0 3 Braund, Mr. Owen Harris male 22.0 1 0 A/5 21171 7.2500 NaN S
1 2 1 1 Cumings, Mrs. John Bradley (Florence Briggs Th... female 38.0 1 0 PC 17599 71.2833 C85 C
2 3 1 3 Heikkinen, Miss. Laina female 26.0 0 0 STON/O2. 3101282 7.9250 NaN S
3 4 1 1 Futrelle, Mrs. Jacques Heath (Lily May Peel) female 35.0 1 0 113803 53.1000 C123 S
4 5 0 3 Allen, Mr. William Henry male 35.0 0 0 373450 8.0500 NaN S
# Let us see how many people survived and who did not.
titanic['Survived'].value_counts()
Survived
0    549
1    342
Name: count, dtype: int64

Basic Python: Variables#

A variable is a name that stores a value. You can use variables to remember numbers, text, or anything else you want to reuse.

Let us try making some variables!

# Store your age and your name
my_age = 15
my_name = "Sam"
print("My name is", my_name)
print("My age is", my_age)
My name is Sam
My age is 15
# Get your favorite color from user input
favorite_color = input("What is your favorite color?")
print("Wow! I like", favorite_color, "too!")
Wow! I like blue too!
 

Selecting Data in Python#

We often want to look at just part of our data.

For example, we might want to see all the women on the Titanic.

This is called filtering or selecting.

# Show all rows where Sex is 'female'
women = titanic[titanic['Sex'] == 'female']
print("Number of women:", len(women))
women.head()
Number of women: 314
PassengerId Survived Pclass Name Sex Age SibSp Parch Ticket Fare Cabin Embarked
1 2 1 1 Cumings, Mrs. John Bradley (Florence Briggs Th... female 38.0 1 0 PC 17599 71.2833 C85 C
2 3 1 3 Heikkinen, Miss. Laina female 26.0 0 0 STON/O2. 3101282 7.9250 NaN S
3 4 1 1 Futrelle, Mrs. Jacques Heath (Lily May Peel) female 35.0 1 0 113803 53.1000 C123 S
8 9 1 3 Johnson, Mrs. Oscar W (Elisabeth Vilhelmina Berg) female 27.0 0 2 347742 11.1333 NaN S
9 10 1 2 Nasser, Mrs. Nicholas (Adele Achem) female 14.0 1 0 237736 30.0708 NaN C

Handling Missing Data#

Real-world data is messy! Sometimes information is missing.

We need to check for missing spots and fill or remove them.

Let us see if our Titanic data has any missing pieces.

# Check for missing values
titanic.isnull().sum()
PassengerId      0
Survived         0
Pclass           0
Name             0
Sex              0
Age            177
SibSp            0
Parch            0
Ticket           0
Fare             0
Cabin          687
Embarked         2
dtype: int64
# Fill missing age data with the average age
mean_age = titanic['Age'].mean()
titanic['Age'].fillna(mean_age, inplace=True)
# Remove columns we do not need
titanic = titanic.drop(['Cabin', 'Ticket', 'Name'], axis=1)

Turning Words into Numbers#

Machine learning models work best with numbers, not words.

We have to turn columns like 'Sex' into numbers.

This process is called encoding.

# Change 'Sex' to a number: 0 for female, 1 for male
titanic['Sex'] = titanic['Sex'].map({'female': 0, 'male': 1})
# Simple model: Predict survival based on Sex
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression

X = titanic[['Sex']]
y = titanic['Survived']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

model = LogisticRegression()
model.fit(X_train, y_train)
LogisticRegression()
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Predict on the test data and check accuracy
y_pred = model.predict(X_test)
accuracy = np.mean(y_pred == y_test)
print("Accuracy:", accuracy)
Accuracy: 0.7821229050279329

Improving the Model#

We can help the model by adding more columns. The more useful information we give it, the better it can learn.

Let us try using age and class along with sex.

# Use more columns: Sex, Age, and Pclass
X = titanic[['Sex', 'Age', 'Pclass']]
y = titanic['Survived']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
model = LogisticRegression()
model.fit(X_train, y_train)
accuracy = model.score(X_test, y_test)
print("New accuracy:", accuracy)
New accuracy: 0.8100558659217877
# Try the model on new data with input().
sex_input = input("Enter 0 for female, 1 for male: ")
age_input = input("Enter age: ")
pclass_input = input("Enter class (1, 2, or 3): ")
# Convert inputs to numbers
user_data = np.array([[int(sex_input), float(age_input), int(pclass_input)]])
prediction = model.predict(user_data)
if prediction[0] == 1:
    print("The model predicts: Survived")
else:
    print("The model predicts: Did not survive")
    
The model predicts: Did not survive
 

Best Practices: Random States and Reproducibility#

A random_state helps us get the same results every time we run.

This is important for sharing your work with others.

You will see random_state used in train_test_split a lot.

# What if we forget random_state?
X_train2, X_test2, y_train2, y_test2 = train_test_split(X, y, test_size=0.2)
print(X_train.equals(X_train2))
False
# Troubleshooting: Common errors
# Try dividing by zero
try:
    result = 10 / 0
except ZeroDivisionError:
    print("Oops! You cannot divide by zero.")
    
Oops! You cannot divide by zero.

Extra Tips#

  • Use comments (the # symbol) to explain your code.
  • Use print statements to check your work.
  • Practice makes perfect. Try changing numbers and running cells again.

You are well on your way to becoming a data scientist!

# Your mini challenge: Try predicting survival using only 'Age' and 'Pclass'
X_challenge = titanic[['Age', 'Pclass']]
y_challenge = titanic['Survived']
X_train_c, X_test_c, y_train_c, y_test_c = train_test_split(X_challenge, y_challenge, test_size=0.2, random_state=42)
model_challenge = LogisticRegression()
model_challenge.fit(X_train_c, y_train_c)
print("Challenge accuracy:", model_challenge.score(X_test_c, y_test_c))
Challenge accuracy: 0.7597765363128491

Capstone Project Recap#

Great job! You loaded real data, cleaned it, built and tested your own model.

You saw how simple steps add up to real data science skills.

With practice, you will improve and try bigger challenges.

Keep experimenting and having fun with your new superpower!

Next Steps and YouTube Call to Action#

Thanks for following along! If you enjoyed this lesson, hit the like button and subscribe for more beginner Python and machine learning projects.

Try changing the data or adding your own twist to the project.

Happy coding!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.