Lesson 26 · Data Science Projects
Predicting Wine Quality with Machine Learning: A Step-by-Step Guide Using Python
This lesson introduces data mining for machine learning using Python. We use the Wine Quality dataset to predict wine quality based on chemical properties.…
- CourseData Science Projects
- Lesson26 of 33
- Video18 min
- FormatJupyter notebook · 15 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWine Quality Prediction with Data Mining#
- This lesson introduces data mining for machine learning using Python.
- We use the Wine Quality dataset to predict wine quality based on chemical properties.
- Data mining helps us find patterns and build models that can predict outcomes.
- Learning these skills is important for anyone interested in real world analytics.
- Let us see how it works step by step!
What is the Wine Quality Dataset?#
- The dataset contains information about red wines.
- Each wine has features like acidity, sugar, pH, and more.
- The target is to predict a quality score between 0 and 10.
- This is a supervised machine learning problem.
# Suppress warnings for a cleaner output
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)
# Data setup
import pandas as pd
url = 'https://archive.ics.uci.edu/ml/machine-learning-databases/wine-quality/winequality-red.csv'
df = pd.read_csv(url, sep=';')
print(df.shape)
print(df.head(3))
Understanding the Data Columns#
- Each row is a red wine sample.
- Features include fixed acidity, volatile acidity, citric acid, sugar, chlorides, and more.
- The last column is 'quality', our target to predict.
# Checking for missing data
df.isnull().sum()
# Basic data statistics
df.describe()
# Show all column names
print(df.columns.tolist())
Data Visualization: Plotting Wine Quality Scores#
- Visualizations help us see patterns at a glance.
- Let us plot how many wines have each quality score.
- It is a quick way to understand the target variable.
# Bar plot of quality scores distribution
import matplotlib.pyplot as plt
df['quality'].value_counts().sort_index().plot(kind='bar', color='skyblue')
plt.xlabel('Wine Quality Score')
plt.ylabel('Count')
plt.title('Wine Quality Score Distribution')
plt.show()
# Visualizing the relationship between alcohol and quality
import seaborn as sns
sns.boxplot(x='quality', y='alcohol', data=df, palette='coolwarm')
plt.title('Alcohol Content vs Quality')
plt.show()
# Correlation heatmap of features
plt.figure(figsize=(10,7))
sns.heatmap(df.corr(), annot=True, fmt='.2f', cmap='YlGnBu')
plt.title('Feature Correlation Heatmap')
plt.show()
Preparing Data for Machine Learning#
- We separate features (X) from the target variable (y).
- Features are the columns used to predict the target.
- The target is the quality score.
# Splitting features and target
X = df.drop('quality', axis=1)
y = df['quality']
# Split data into train and test sets
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print(X_train.shape, X_test.shape)
Model Building: Random Forest Classifier#
- Random Forest is a popular machine learning algorithm for classification.
- It builds many decision trees and combines their results.
- We use it to predict wine quality with the given features.
# Training the model
from sklearn.ensemble import RandomForestClassifier
model = RandomForestClassifier(random_state=42)
model.fit(X_train, y_train)
# Making predictions on the test set
y_pred = model.predict(X_test)
print(y_pred[:10])
# Evaluating our model
from sklearn.metrics import accuracy_score, classification_report
acc = accuracy_score(y_test, y_pred)
print(f'Accuracy: {acc:.2f}')
print(classification_report(y_test, y_pred))
# Feature importance plot
import numpy as np
importances = model.feature_importances_
indices = np.argsort(importances)[::-1]
plt.figure(figsize=(8,5))
sns.barplot(x=importances[indices], y=X.columns[indices], palette='viridis')
plt.title('Feature Importance for Predicting Wine Quality')
plt.xlabel('Importance')
plt.ylabel('Feature')
plt.show()
# Interactive: Predict wine quality for your own sample
sample = []
for col in X.columns:
val = float(input(f'Enter value for {col}: '))
sample.append(val)
sample = np.array(sample).reshape(1, -1)
pred = model.predict(sample)
print(f'Predicted quality: {pred[0]}')
Summary: What We Learned#
- We loaded the Wine Quality dataset and explored it visually.
- We checked for missing data and looked at feature statistics.
- We trained a Random Forest model to predict quality.
- We measured performance and checked which features are important.
- Changing feature values can change wine quality predictions.
- This is a foundation for many kinds of data mining problems.
- If you enjoyed this lesson, like and subscribe for more!
Next Steps#
- Try other machine learning algorithms like Support Vector Machines or Gradient Boosting.
- See if scaling the features improves the model performance.
- You can use the same approach with your own datasets.
- Leave a comment with your results!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



