Mathew K Analytics

Lesson 51 · Data Science Projects

Building a Neural Network for Handwritten Digit Recognition: A Step-by-Step Guide

In this beginner friendly lesson, you will learn to recognize handwritten digits using Python and deep learning. We will use the famous MNIST dataset of…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Handwritten Digit Recognition with Neural Networks#

  • In this beginner friendly lesson, you will learn to recognize handwritten digits using Python and deep learning.
  • We will use the famous MNIST dataset of digit images.
  • You will explore how neural networks can convert images into predictions.
  • This skill is important for computer vision and modern AI applications.
  • By the end, you will be able to build, train, and test a simple neural network model.
# Suppress warnings for a clean output
import warnings; warnings.filterwarnings("ignore")
import numpy as np
np.random.seed(42)

What is the MNIST Dataset?#

  • The MNIST dataset is a large set of handwritten digits.
  • It contains 70000 images of numbers from 0 to 9.
  • Each image is 28 by 28 pixels and comes with the actual digit as a label.
  • Data scientists use MNIST to test computer vision and neural network models easily.
# Data setup
from tensorflow.keras.datasets import mnist
(X_train, y_train), (X_test, y_test) = mnist.load_data()
print(X_train.shape, y_train.shape)
print(X_test.shape, y_test.shape)
(60000, 28, 28) (60000,)
(10000, 28, 28) (10000,)
# Let us preview an image and its label
import matplotlib.pyplot as plt
plt.imshow(X_train[0], cmap="gray")
plt.title(f"Label: {y_train[0]}")
plt.axis("off")
plt.show()
No description has been provided for this image

Why Do We Need Neural Networks?#

  • Traditional computer programs cannot easily tell digits apart from pixels.
  • Neural networks learn from the raw image data and improve with practice.
  • They are inspired by how human brains process patterns in vision.
# Data normalization for better learning
X_train = X_train / 255.0
X_test = X_test / 255.0
# Flatten the 28x28 images into 1D vectors
X_train_flat = X_train.reshape(X_train.shape[0], -1)
X_test_flat = X_test.reshape(X_test.shape[0], -1)
# Let us confirm the new shape
print(X_train_flat.shape)
print(X_test_flat.shape)
(60000, 784)
(10000, 784)

What is a Neural Network?#

  • A neural network is a set of connected layers where each layer transforms the input data.
  • It learns patterns by adjusting weights using data and feedback.
  • In this lesson, we will use a simple neural network with a few layers.
# Build a simple neural network model
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense
model = Sequential([
    Dense(128, activation="relu", input_shape=(784,)),
    Dense(64, activation="relu"),
    Dense(10, activation="softmax")
])
model.summary()
Model: "sequential"
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Layer (type)                    ┃ Output Shape           ┃       Param # ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ dense (Dense)                   │ (None, 128)            │       100,480 │
├─────────────────────────────────┼────────────────────────┼───────────────┤
│ dense_1 (Dense)                 │ (None, 64)             │         8,256 │
├─────────────────────────────────┼────────────────────────┼───────────────┤
│ dense_2 (Dense)                 │ (None, 10)             │           650 │
└─────────────────────────────────┴────────────────────────┴───────────────┘
 Total params: 109,386 (427.29 KB)
 Trainable params: 109,386 (427.29 KB)
 Non-trainable params: 0 (0.00 B)
# Compile the model with loss, optimizer, and metrics
model.compile(
    loss="sparse_categorical_crossentropy",
    optimizer="adam",
    metrics=["accuracy"]
)
# Train the neural network
history = model.fit(X_train_flat, y_train, epochs=5, batch_size=32, validation_split=0.1)
Epoch 1/5
1688/1688 ━━━━━━━━━━━━━━━━━━━━ 5s 2ms/step - accuracy: 0.9253 - loss: 0.2508 - val_accuracy: 0.9670 - val_loss: 0.1107
Epoch 2/5
1688/1688 ━━━━━━━━━━━━━━━━━━━━ 4s 2ms/step - accuracy: 0.9679 - loss: 0.1057 - val_accuracy: 0.9723 - val_loss: 0.0929
Epoch 3/5
1688/1688 ━━━━━━━━━━━━━━━━━━━━ 4s 2ms/step - accuracy: 0.9780 - loss: 0.0713 - val_accuracy: 0.9692 - val_loss: 0.0985
Epoch 4/5
1688/1688 ━━━━━━━━━━━━━━━━━━━━ 4s 2ms/step - accuracy: 0.9821 - loss: 0.0564 - val_accuracy: 0.9767 - val_loss: 0.0759
Epoch 5/5
1688/1688 ━━━━━━━━━━━━━━━━━━━━ 4s 2ms/step - accuracy: 0.9859 - loss: 0.0436 - val_accuracy: 0.9765 - val_loss: 0.0971
# Plot training and validation accuracy
import matplotlib.pyplot as plt
plt.plot(history.history["accuracy"], label="Training Accuracy")
plt.plot(history.history["val_accuracy"], label="Validation Accuracy")
plt.xlabel("Epoch")
plt.ylabel("Accuracy")
plt.title("Model Training Progress")
plt.legend()
plt.show()
No description has been provided for this image
# Evaluate on test data to check real-world performance
test_loss, test_acc = model.evaluate(X_test_flat, y_test)
print(f"Test accuracy: {test_acc:.2f}")
313/313 ━━━━━━━━━━━━━━━━━━━━ 1s 2ms/step - accuracy: 0.9740 - loss: 0.0893
Test accuracy: 0.97
# Predict a digit and show the result
import numpy as np
idx = np.random.randint(0, X_test_flat.shape[0])
image = X_test_flat[idx]
label = y_test[idx]
pred = model.predict(image.reshape(1, -1))
predicted_digit = np.argmax(pred)
plt.imshow(X_test[idx], cmap="gray")
plt.title(f"Predicted: {predicted_digit}, True: {label}")
plt.axis("off")
plt.show()
1/1 ━━━━━━━━━━━━━━━━━━━━ 0s 111ms/step
No description has been provided for this image
# Make multiple predictions and check overall performance
preds = model.predict(X_test_flat)
predicted_labels = np.argmax(preds, axis=1)
from sklearn.metrics import classification_report, confusion_matrix
print(classification_report(y_test, predicted_labels))
313/313 ━━━━━━━━━━━━━━━━━━━━ 0s 1ms/step
              precision    recall  f1-score   support

           0       0.99      0.98      0.98       980
           1       0.99      0.98      0.99      1135
           2       0.97      0.98      0.97      1032
           3       0.97      0.98      0.97      1010
           4       0.96      0.98      0.97       982
           5       0.99      0.95      0.97       892
           6       0.97      0.98      0.98       958
           7       0.98      0.97      0.98      1028
           8       0.97      0.96      0.96       974
           9       0.95      0.98      0.96      1009

    accuracy                           0.97     10000
   macro avg       0.97      0.97      0.97     10000
weighted avg       0.97      0.97      0.97     10000

# Show a confusion matrix of real vs predicted digits
import seaborn as sns
cm = confusion_matrix(y_test, predicted_labels)
plt.figure(figsize=(8,6))
sns.heatmap(cm, annot=True, fmt="d", cmap="Blues")
plt.xlabel("Predicted Digit")
plt.ylabel("True Digit")
plt.title("Confusion Matrix")
plt.show()
No description has been provided for this image

Recap: What Have You Achieved?#

  • You loaded and explored a real world image dataset.
  • You built and trained a working neural network from scratch.
  • You evaluated its skill at reading handwritten numbers.
  • You visualized where it succeeds and where it makes mistakes.
  • You are now ready to explore deeper networks or new datasets!
# Mini project challenge: Try with fewer data or more epochs
choice = input("Type '1' to use half the data, or '2' to train for 10 epochs: ")
if choice == '1':
    print("Training with half the data...")
    subset_X = X_train_flat[:30000]
    subset_y = y_train[:30000]
    model.fit(subset_X, subset_y, epochs=5, batch_size=32, validation_split=0.1)
elif choice == '2':
    print("Training for 10 epochs...")
    model.fit(X_train_flat, y_train, epochs=10, batch_size=32, validation_split=0.1)
else:
    print("Try rerunning and enter 1 or 2.")
Training with half the data...
Epoch 1/5
844/844 ━━━━━━━━━━━━━━━━━━━━ 2s 3ms/step - accuracy: 0.9903 - loss: 0.0283 - val_accuracy: 0.9873 - val_loss: 0.0346
Epoch 2/5
844/844 ━━━━━━━━━━━━━━━━━━━━ 2s 3ms/step - accuracy: 0.9933 - loss: 0.0191 - val_accuracy: 0.9890 - val_loss: 0.0321
Epoch 3/5
844/844 ━━━━━━━━━━━━━━━━━━━━ 2s 3ms/step - accuracy: 0.9938 - loss: 0.0187 - val_accuracy: 0.9883 - val_loss: 0.0389
Epoch 4/5
844/844 ━━━━━━━━━━━━━━━━━━━━ 2s 3ms/step - accuracy: 0.9956 - loss: 0.0138 - val_accuracy: 0.9857 - val_loss: 0.0523
Epoch 5/5
844/844 ━━━━━━━━━━━━━━━━━━━━ 2s 3ms/step - accuracy: 0.9946 - loss: 0.0158 - val_accuracy: 0.9850 - val_loss: 0.0489

Next Steps and Extra Ideas#

  • Experiment with deeper or wider networks to see if accuracy improves.
  • Try using image data directly without flattening, by adding Conv2D layers for more advanced learning.
  • Make a website or app where users upload their own digit drawings for instant recognition.
  • Share your cool results with friends! Did your accuracy change with different options?

Thank You for Learning With Us#

  • If you enjoyed this lesson, subscribe for more beginner friendly tutorials.
  • Like, comment, and share if you want more deep learning demos!
  • You are now confident with neural networks for digit recognition.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.