Lesson 83 · Data Science Projects
Detecting Fake News with NLP and Transformer Models: A Biblical Perspective
In this lesson, we will explore how to detect fake news using Natural Language Processing (NLP) and Transformers like BERT. We will analyze text data, clean…
- CourseData Science Projects
- Lesson83 of 33
- Video15 min
- FormatJupyter notebook · 14 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWeek 9: Fake News Detection with NLP and Transformers#
In this lesson, we will explore how to detect fake news using Natural Language Processing (NLP) and Transformers like BERT.
We will analyze text data, clean it, transform it into features the computer can understand, and build a machine learning model.
This skill is used by social media, news platforms, and fact-checking organizations to spot misleading information.
# Suppress warnings to keep output clean
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Data setup (Fake News Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/lutzhamel/fake-news/master/data/fake_or_real_news.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
What does the data look like?#
Each row is an online news article headline.
There is a 'label' column showing if it is 'FAKE' or 'REAL'.
# Check for missing values
print(df.isnull().sum())
# Clean up missing or bad data (if any)
df = df.dropna()
print(df.shape)
# Look at the balance of real vs fake news
print(df['label'].value_counts())
Practice Prompt#
Can you think of two reasons fake news is a challenge for society?
Pause and write them down!
# Explore an article headline
print(df.iloc[0]['text'])
# Convert labels to numbers: REAL=1, FAKE=0
df['label_num'] = df['label'].map({'REAL': 1, 'FAKE': 0})
# Split data into training and test sets
from sklearn.model_selection import train_test_split
X = df['text']
y = df['label_num']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
What is NLP?#
NLP stands for Natural Language Processing.
It is a field of AI that helps computers understand and work with human language.
# Convert text to features with TF-IDF
from sklearn.feature_extraction.text import TfidfVectorizer
tfidf = TfidfVectorizer(stop_words='english', max_df=0.7)
X_train_tfidf = tfidf.fit_transform(X_train)
X_test_tfidf = tfidf.transform(X_test)
# Simple model: Logistic Regression
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=1000)
model.fit(X_train_tfidf, y_train)
# Test accuracy
score = model.score(X_test_tfidf, y_test)
print(f"Test accuracy: {score:.2f}")
# Make a fake news prediction from user input
headline = input("Type a news headline to check: ")
features = tfidf.transform([headline])
pred = model.predict(features)[0]
if pred == 1:
print("This article looks REAL.")
else:
print("This article looks FAKE.")
How do Transformers work?#
Transformers like BERT can understand language much more deeply than older models.
They read the whole sentence and capture the real meaning of words as they appear in context.
This is helpful in detecting sarcasm, lies, and subtle tricks in fake news.
# Use BERT transformer with Hugging Face to make predictions
from transformers import pipeline
bert_clf = pipeline('text-classification', model='distilbert-base-uncased-finetuned-sst-2-english')
result = bert_clf("Breaking news: The moon is filled with cheese!")
print(result)
Practice: Try your own headline in BERT!#
Run the cell above again, but change the text to a real or fake news headline you make up.
Notice how it predicts the label.
# Visualize word importance (feature weights)
import numpy as np
import matplotlib.pyplot as plt
feature_names = np.array(tfidf.get_feature_names_out())
coefs = model.coef_[0]
top_features = 10
top_pos = np.argsort(coefs)[-top_features:]
top_neg = np.argsort(coefs)[:top_features]
plt.figure(figsize=(10, 4))
plt.barh(feature_names[top_pos], coefs[top_pos], color='green')
plt.barh(feature_names[top_neg], coefs[top_neg], color='red')
plt.xlabel('Feature weight')
plt.title('Words most linked to REAL (green) and FAKE (red) articles')
plt.show()
Recap#
You learned how to:
Load and explore the Fake News dataset Clean text data and encode labels Split data into training and test sets Turn headlines into features using TF-IDF Train a logistic regression model Test the model and use it to predict new headlines Try BERT transformer for deeper understanding Visualize top words for fake vs real news
Congratulations! You took your first steps towards building practical NLP tools.
Next Steps#
Try changing the classifier to Random Forest or test new headlines.
Check out other lessons for more hands-on projects!
If you learned something new, subscribe on YouTube and keep practicing!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



