Lesson 15 · Data Science Projects
Analyzing and Predicting Iris Flower Species Using Data Science Techniques
Discovering patterns in famous flower data Classification, visualization, and real world insights See how analysis leads to predictions No coding experience…
- CourseData Science Projects
- Lesson15 of 33
- Video22 min
- FormatJupyter notebook · 16 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIris Flower Dataset Analysis and Predictions#
- Discovering patterns in famous flower data
- Classification, visualization, and real world insights
- See how analysis leads to predictions
- No coding experience needed
What is the Iris Dataset?#
- A classic dataset for data science
- Measures four features of three flower species
- Used to learn classification and data mining
- Easy to visualize and great for beginners
# Suppress warnings for a clean output
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# Data setup
import pandas as pd
url = 'https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
# What columns are in our dataset?
print(df.columns.tolist())
# Basic info and missing value check
info = df.info()
missing = df.isnull().sum()
print(missing)
# Show summary statistics
print(df.describe())
Visualizing the Data#
- Pictures reveal hidden patterns fast
- We can quickly spot relationships and trends
- Visualization helps us choose analysis paths
# Plotting petal length vs petal width
import matplotlib.pyplot as plt
plt.figure(figsize=(6,4))
for species in df['species'].unique():
plt.scatter(df[df['species'] == species]['petal_length'],
df[df['species'] == species]['petal_width'], label=species)
plt.xlabel('Petal Length')
plt.ylabel('Petal Width')
plt.title('Petal Size by Flower Species')
plt.legend()
plt.show()
# Histogram of sepal length for all flowers
df['sepal_length'].hist(bins=20, color='skyblue', edgecolor='black')
plt.xlabel('Sepal Length')
plt.ylabel('Frequency')
plt.title('Distribution of Sepal Lengths in Dataset')
plt.show()
# Pairplot for relationships among all features
import seaborn as sns
sns.pairplot(df, hue='species', diag_kind='hist')
plt.show()
Preparing the Data for Prediction#
- Machine learning needs numeric inputs only
- We will split our data into train and test sets
- We use the train data to learn patterns
- Test data lets us check how well our future predictions might do
# Convert species names to numbers
from sklearn.preprocessing import LabelEncoder
le = LabelEncoder()
df['species_num'] = le.fit_transform(df['species'])
print(df[['species', 'species_num']].head(6))
# Split data into train and test sets (80-20 split)
from sklearn.model_selection import train_test_split
features = df.drop(['species', 'species_num'], axis=1).columns
X = df[features]
y = df['species_num']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print('Train:', X_train.shape, 'Test:', X_test.shape)
Predicting Species with K-Nearest Neighbors#
- A simple model that checks which flowers are closest
- If neighbors agree, predicts their type
- KNN is popular for beginners and works well here
# Fit a KNN classifier
from sklearn.neighbors import KNeighborsClassifier
knn = KNeighborsClassifier(n_neighbors=5)
knn.fit(X_train, y_train)
score = knn.score(X_test, y_test)
print(f"KNN accuracy: {score:.2%}")
# What features matter most? Try with fewer features
X_train_small = X_train[['petal_length', 'petal_width']]
X_test_small = X_test[['petal_length', 'petal_width']]
knn2 = KNeighborsClassifier(n_neighbors=5)
knn2.fit(X_train_small, y_train)
score2 = knn2.score(X_test_small, y_test)
print(f"KNN with petals only: {score2:.2%}")
# Predict the class of a new flower (interactive)
s_length = float(input("Enter sepal length in cm: "))
s_width = float(input("Enter sepal width in cm: "))
p_length = float(input("Enter petal length in cm: "))
p_width = float(input("Enter petal width in cm: "))
sample = [[s_length, s_width, p_length, p_width]]
pred = knn.predict(sample)[0]
print(f"Predicted species: {le.inverse_transform([pred])[0]}")
Try It Yourself: Mini Exercise#
- Use the interactive prediction cell again
- Try flower measurements from each species group
- Can you 'trick' the model?
- What happens if you enter values out of the shown ranges?
# Confusion matrix for detailed performance
from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay
import numpy as np
y_pred = knn.predict(X_test)
cm = confusion_matrix(y_test, y_pred)
disp = ConfusionMatrixDisplay(confusion_matrix=cm, display_labels=le.classes_)
disp.plot(cmap='Blues')
plt.title('KNN Classification Confusion Matrix')
plt.show()
# Can other algorithms do better? Try a decision tree
from sklearn.tree import DecisionTreeClassifier
tree = DecisionTreeClassifier(random_state=42)
tree.fit(X_train, y_train)
tree_score = tree.score(X_test, y_test)
print(f"Decision Tree accuracy: {tree_score:.2%}")
# Visualize the actual tree structure (simplified)
from sklearn import tree as sktree
plt.figure(figsize=(10,6))
sktree.plot_tree(tree, feature_names=features, class_names=le.classes_, filled=True, max_depth=2)
plt.title('Decision Tree Structure (First 2 Levels)')
plt.show()
Next Steps and Practice#
- Try other classifiers on this dataset (Random Forest, Logistic Regression)
- Practice more plotting with seaborn or matplotlib
- Can you create your own small flower dataset and predict manually?
- If you liked this video, subscribe and comment below for more beginner friendly data lessons!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



