Mathew K Analytics

Lesson 15 · Data Science Projects

Analyzing and Predicting Iris Flower Species Using Data Science Techniques

Discovering patterns in famous flower data Classification, visualization, and real world insights See how analysis leads to predictions No coding experience…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Iris Flower Dataset Analysis and Predictions#

  • Discovering patterns in famous flower data
  • Classification, visualization, and real world insights
  • See how analysis leads to predictions
  • No coding experience needed

What is the Iris Dataset?#

  • A classic dataset for data science
  • Measures four features of three flower species
  • Used to learn classification and data mining
  • Easy to visualize and great for beginners
# Suppress warnings for a clean output
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# Data setup
import pandas as pd
url = 'https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(150, 5)
   sepal_length  sepal_width  petal_length  petal_width species
0           5.1          3.5           1.4          0.2  setosa
1           4.9          3.0           1.4          0.2  setosa
2           4.7          3.2           1.3          0.2  setosa
# What columns are in our dataset?
print(df.columns.tolist())
['sepal_length', 'sepal_width', 'petal_length', 'petal_width', 'species']
# Basic info and missing value check
info = df.info()
missing = df.isnull().sum()
print(missing)
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 150 entries, 0 to 149
Data columns (total 5 columns):
 #   Column        Non-Null Count  Dtype  
---  ------        --------------  -----  
 0   sepal_length  150 non-null    float64
 1   sepal_width   150 non-null    float64
 2   petal_length  150 non-null    float64
 3   petal_width   150 non-null    float64
 4   species       150 non-null    object 
dtypes: float64(4), object(1)
memory usage: 6.0+ KB
sepal_length    0
sepal_width     0
petal_length    0
petal_width     0
species         0
dtype: int64
# Show summary statistics
print(df.describe())
       sepal_length  sepal_width  petal_length  petal_width
count    150.000000   150.000000    150.000000   150.000000
mean       5.843333     3.054000      3.758667     1.198667
std        0.828066     0.433594      1.764420     0.763161
min        4.300000     2.000000      1.000000     0.100000
25%        5.100000     2.800000      1.600000     0.300000
50%        5.800000     3.000000      4.350000     1.300000
75%        6.400000     3.300000      5.100000     1.800000
max        7.900000     4.400000      6.900000     2.500000

Visualizing the Data#

  • Pictures reveal hidden patterns fast
  • We can quickly spot relationships and trends
  • Visualization helps us choose analysis paths
# Plotting petal length vs petal width
import matplotlib.pyplot as plt
plt.figure(figsize=(6,4))
for species in df['species'].unique():
    plt.scatter(df[df['species'] == species]['petal_length'],
                df[df['species'] == species]['petal_width'], label=species)
plt.xlabel('Petal Length')
plt.ylabel('Petal Width')
plt.title('Petal Size by Flower Species')
plt.legend()
plt.show()
No description has been provided for this image
# Histogram of sepal length for all flowers
df['sepal_length'].hist(bins=20, color='skyblue', edgecolor='black')
plt.xlabel('Sepal Length')
plt.ylabel('Frequency')
plt.title('Distribution of Sepal Lengths in Dataset')
plt.show()
No description has been provided for this image
# Pairplot for relationships among all features
import seaborn as sns
sns.pairplot(df, hue='species', diag_kind='hist')
plt.show()
No description has been provided for this image

Preparing the Data for Prediction#

  • Machine learning needs numeric inputs only
  • We will split our data into train and test sets
  • We use the train data to learn patterns
  • Test data lets us check how well our future predictions might do
# Convert species names to numbers
from sklearn.preprocessing import LabelEncoder
le = LabelEncoder()
df['species_num'] = le.fit_transform(df['species'])
print(df[['species', 'species_num']].head(6))
  species  species_num
0  setosa            0
1  setosa            0
2  setosa            0
3  setosa            0
4  setosa            0
5  setosa            0
# Split data into train and test sets (80-20 split)
from sklearn.model_selection import train_test_split
features = df.drop(['species', 'species_num'], axis=1).columns
X = df[features]
y = df['species_num']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print('Train:', X_train.shape, 'Test:', X_test.shape)
Train: (120, 4) Test: (30, 4)

Predicting Species with K-Nearest Neighbors#

  • A simple model that checks which flowers are closest
  • If neighbors agree, predicts their type
  • KNN is popular for beginners and works well here
# Fit a KNN classifier
from sklearn.neighbors import KNeighborsClassifier
knn = KNeighborsClassifier(n_neighbors=5)
knn.fit(X_train, y_train)
score = knn.score(X_test, y_test)
print(f"KNN accuracy: {score:.2%}")
KNN accuracy: 100.00%
# What features matter most? Try with fewer features
X_train_small = X_train[['petal_length', 'petal_width']]
X_test_small = X_test[['petal_length', 'petal_width']]
knn2 = KNeighborsClassifier(n_neighbors=5)
knn2.fit(X_train_small, y_train)
score2 = knn2.score(X_test_small, y_test)
print(f"KNN with petals only: {score2:.2%}")
KNN with petals only: 100.00%
# Predict the class of a new flower (interactive)
s_length = float(input("Enter sepal length in cm: "))
s_width = float(input("Enter sepal width in cm: "))
p_length = float(input("Enter petal length in cm: "))
p_width = float(input("Enter petal width in cm: "))
sample = [[s_length, s_width, p_length, p_width]]
pred = knn.predict(sample)[0]
print(f"Predicted species: {le.inverse_transform([pred])[0]}")
Predicted species: setosa

Try It Yourself: Mini Exercise#

  • Use the interactive prediction cell again
  • Try flower measurements from each species group
  • Can you 'trick' the model?
  • What happens if you enter values out of the shown ranges?
# Confusion matrix for detailed performance
from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay
import numpy as np
y_pred = knn.predict(X_test)
cm = confusion_matrix(y_test, y_pred)
disp = ConfusionMatrixDisplay(confusion_matrix=cm, display_labels=le.classes_)
disp.plot(cmap='Blues')
plt.title('KNN Classification Confusion Matrix')
plt.show()
No description has been provided for this image
# Can other algorithms do better? Try a decision tree
from sklearn.tree import DecisionTreeClassifier
tree = DecisionTreeClassifier(random_state=42)
tree.fit(X_train, y_train)
tree_score = tree.score(X_test, y_test)
print(f"Decision Tree accuracy: {tree_score:.2%}")
Decision Tree accuracy: 100.00%
# Visualize the actual tree structure (simplified)
from sklearn import tree as sktree
plt.figure(figsize=(10,6))
sktree.plot_tree(tree, feature_names=features, class_names=le.classes_, filled=True, max_depth=2)
plt.title('Decision Tree Structure (First 2 Levels)')
plt.show()
No description has been provided for this image

Next Steps and Practice#

  • Try other classifiers on this dataset (Random Forest, Logistic Regression)
  • Practice more plotting with seaborn or matplotlib
  • Can you create your own small flower dataset and predict manually?
  • If you liked this video, subscribe and comment below for more beginner friendly data lessons!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.