Mathew K Analytics

Lesson 26 · Data Science Projects

Predicting Wine Quality with Machine Learning: A Step-by-Step Guide Using Python

This lesson introduces data mining for machine learning using Python. We use the Wine Quality dataset to predict wine quality based on chemical properties.…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Wine Quality Prediction with Data Mining#

  • This lesson introduces data mining for machine learning using Python.
  • We use the Wine Quality dataset to predict wine quality based on chemical properties.
  • Data mining helps us find patterns and build models that can predict outcomes.
  • Learning these skills is important for anyone interested in real world analytics.
  • Let us see how it works step by step!

What is the Wine Quality Dataset?#

  • The dataset contains information about red wines.
  • Each wine has features like acidity, sugar, pH, and more.
  • The target is to predict a quality score between 0 and 10.
  • This is a supervised machine learning problem.
# Suppress warnings for a cleaner output
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
 
# Data setup
import pandas as pd
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/wine-quality/winequality-red.csv'
df = pd.read_csv(url, sep=';')
print(df.shape)
print(df.head(3))
(1599, 12)
   fixed acidity  volatile acidity  citric acid  residual sugar  chlorides  \
0            7.4              0.70         0.00             1.9      0.076   
1            7.8              0.88         0.00             2.6      0.098   
2            7.8              0.76         0.04             2.3      0.092   

   free sulfur dioxide  total sulfur dioxide  density    pH  sulphates  \
0                 11.0                  34.0   0.9978  3.51       0.56   
1                 25.0                  67.0   0.9968  3.20       0.68   
2                 15.0                  54.0   0.9970  3.26       0.65   

   alcohol  quality  
0      9.4        5  
1      9.8        5  
2      9.8        5  

Understanding the Data Columns#

  • Each row is a red wine sample.
  • Features include fixed acidity, volatile acidity, citric acid, sugar, chlorides, and more.
  • The last column is 'quality', our target to predict.
# Checking for missing data
df.isnull().sum()
fixed acidity           0
volatile acidity        0
citric acid             0
residual sugar          0
chlorides               0
free sulfur dioxide     0
total sulfur dioxide    0
density                 0
pH                      0
sulphates               0
alcohol                 0
quality                 0
dtype: int64
# Basic data statistics
df.describe()
fixed acidity volatile acidity citric acid residual sugar chlorides free sulfur dioxide total sulfur dioxide density pH sulphates alcohol quality
count 1599.000000 1599.000000 1599.000000 1599.000000 1599.000000 1599.000000 1599.000000 1599.000000 1599.000000 1599.000000 1599.000000 1599.000000
mean 8.319637 0.527821 0.270976 2.538806 0.087467 15.874922 46.467792 0.996747 3.311113 0.658149 10.422983 5.636023
std 1.741096 0.179060 0.194801 1.409928 0.047065 10.460157 32.895324 0.001887 0.154386 0.169507 1.065668 0.807569
min 4.600000 0.120000 0.000000 0.900000 0.012000 1.000000 6.000000 0.990070 2.740000 0.330000 8.400000 3.000000
25% 7.100000 0.390000 0.090000 1.900000 0.070000 7.000000 22.000000 0.995600 3.210000 0.550000 9.500000 5.000000
50% 7.900000 0.520000 0.260000 2.200000 0.079000 14.000000 38.000000 0.996750 3.310000 0.620000 10.200000 6.000000
75% 9.200000 0.640000 0.420000 2.600000 0.090000 21.000000 62.000000 0.997835 3.400000 0.730000 11.100000 6.000000
max 15.900000 1.580000 1.000000 15.500000 0.611000 72.000000 289.000000 1.003690 4.010000 2.000000 14.900000 8.000000
# Show all column names
print(df.columns.tolist())
['fixed acidity', 'volatile acidity', 'citric acid', 'residual sugar', 'chlorides', 'free sulfur dioxide', 'total sulfur dioxide', 'density', 'pH', 'sulphates', 'alcohol', 'quality']

Data Visualization: Plotting Wine Quality Scores#

  • Visualizations help us see patterns at a glance.
  • Let us plot how many wines have each quality score.
  • It is a quick way to understand the target variable.
# Bar plot of quality scores distribution
import matplotlib.pyplot as plt
df['quality'].value_counts().sort_index().plot(kind='bar', color='skyblue')
plt.xlabel('Wine Quality Score')
plt.ylabel('Count')
plt.title('Wine Quality Score Distribution')
plt.show()
No description has been provided for this image
# Visualizing the relationship between alcohol and quality
import seaborn as sns
sns.boxplot(x='quality', y='alcohol', data=df, palette='coolwarm')
plt.title('Alcohol Content vs Quality')
plt.show()
No description has been provided for this image
# Correlation heatmap of features
plt.figure(figsize=(10,7))
sns.heatmap(df.corr(), annot=True, fmt='.2f', cmap='YlGnBu')
plt.title('Feature Correlation Heatmap')
plt.show()
No description has been provided for this image

Preparing Data for Machine Learning#

  • We separate features (X) from the target variable (y).
  • Features are the columns used to predict the target.
  • The target is the quality score.
# Splitting features and target
X = df.drop('quality', axis=1)
y = df['quality']
# Split data into train and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print(X_train.shape, X_test.shape)
(1279, 11) (320, 11)

Model Building: Random Forest Classifier#

  • Random Forest is a popular machine learning algorithm for classification.
  • It builds many decision trees and combines their results.
  • We use it to predict wine quality with the given features.
# Training the model
from sklearn.ensemble import RandomForestClassifier
model = RandomForestClassifier(random_state=42)
model.fit(X_train, y_train)
RandomForestClassifier(random_state=42)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Making predictions on the test set
y_pred = model.predict(X_test)
print(y_pred[:10])
[5 5 5 5 6 5 5 5 6 6]
# Evaluating our model
from sklearn.metrics import accuracy_score, classification_report
acc = accuracy_score(y_test, y_pred)
print(f'Accuracy: {acc:.2f}')
print(classification_report(y_test, y_pred))
Accuracy: 0.66
              precision    recall  f1-score   support

           3       0.00      0.00      0.00         1
           4       0.00      0.00      0.00        10
           5       0.72      0.75      0.73       130
           6       0.63      0.69      0.66       132
           7       0.63      0.52      0.57        42
           8       0.00      0.00      0.00         5

    accuracy                           0.66       320
   macro avg       0.33      0.33      0.33       320
weighted avg       0.63      0.66      0.64       320

# Feature importance plot
import numpy as np
importances = model.feature_importances_
indices = np.argsort(importances)[::-1]
plt.figure(figsize=(8,5))
sns.barplot(x=importances[indices], y=X.columns[indices], palette='viridis')
plt.title('Feature Importance for Predicting Wine Quality')
plt.xlabel('Importance')
plt.ylabel('Feature')
plt.show()
No description has been provided for this image
# Interactive: Predict wine quality for your own sample
sample = []
for col in X.columns:
    val = float(input(f'Enter value for {col}: '))
    sample.append(val)
sample = np.array(sample).reshape(1, -1)
pred = model.predict(sample)
print(f'Predicted quality: {pred[0]}')
Predicted quality: 5

Summary: What We Learned#

  • We loaded the Wine Quality dataset and explored it visually.
  • We checked for missing data and looked at feature statistics.
  • We trained a Random Forest model to predict quality.
  • We measured performance and checked which features are important.
  • Changing feature values can change wine quality predictions.
  • This is a foundation for many kinds of data mining problems.
  • If you enjoyed this lesson, like and subscribe for more!

Next Steps#

  • Try other machine learning algorithms like Support Vector Machines or Gradient Boosting.
  • See if scaling the features improves the model performance.
  • You can use the same approach with your own datasets.
  • Leave a comment with your results!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.