Mathew K Analytics

Lesson 83 · Data Science Projects

Detecting Fake News with NLP and Transformer Models: A Biblical Perspective

In this lesson, we will explore how to detect fake news using Natural Language Processing (NLP) and Transformers like BERT. We will analyze text data, clean…

⬇ Download notebookOpen in Colab ↗
scikit-learnNumPypandasHugging Face TransformersMatplotlib

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Week 9: Fake News Detection with NLP and Transformers#

In this lesson, we will explore how to detect fake news using Natural Language Processing (NLP) and Transformers like BERT.

We will analyze text data, clean it, transform it into features the computer can understand, and build a machine learning model.

This skill is used by social media, news platforms, and fact-checking organizations to spot misleading information.

# Suppress warnings to keep output clean
import warnings
import numpy as np
np.random.seed(42)
warnings.filterwarnings("ignore")
# Data setup (Fake News Dataset)
import pandas as pd
url = 'https://raw.githubusercontent.com/lutzhamel/fake-news/master/data/fake_or_real_news.csv'
df = pd.read_csv(url)
print(df.shape)
print(df.head(3))
(6335, 4)
      id                                              title  \
0   8476                       You Can Smell Hillary’s Fear   
1  10294  Watch The Exact Moment Paul Ryan Committed Pol...   
2   3608        Kerry to go to Paris in gesture of sympathy   

                                                text label  
0  Daniel Greenfield, a Shillman Journalism Fello...  FAKE  
1  Google Pinterest Digg Linkedin Reddit Stumbleu...  FAKE  
2  U.S. Secretary of State John F. Kerry said Mon...  REAL  

What does the data look like?#

Each row is an online news article headline.

There is a 'label' column showing if it is 'FAKE' or 'REAL'.

# Check for missing values
print(df.isnull().sum())
id       0
title    0
text     0
label    0
dtype: int64
# Clean up missing or bad data (if any)
df = df.dropna()
print(df.shape)
(6335, 4)
# Look at the balance of real vs fake news
print(df['label'].value_counts())
label
REAL    3171
FAKE    3164
Name: count, dtype: int64

Practice Prompt#

Can you think of two reasons fake news is a challenge for society?

Pause and write them down!

# Explore an article headline
print(df.iloc[0]['text'])
Daniel Greenfield, a Shillman Journalism Fellow at the Freedom Center, is a New York writer focusing on radical Islam. 
In the final stretch of the election, Hillary Rodham Clinton has gone to war with the FBI. 
The word “unprecedented” has been thrown around so often this election that it ought to be retired. But it’s still unprecedented for the nominee of a major political party to go war with the FBI. 
But that’s exactly what Hillary and her people have done. Coma patients just waking up now and watching an hour of CNN from their hospital beds would assume that FBI Director James Comey is Hillary’s opponent in this election. 
The FBI is under attack by everyone from Obama to CNN. Hillary’s people have circulated a letter attacking Comey. There are currently more media hit pieces lambasting him than targeting Trump. It wouldn’t be too surprising if the Clintons or their allies were to start running attack ads against the FBI. 
The FBI’s leadership is being warned that the entire left-wing establishment will form a lynch mob if they continue going after Hillary. And the FBI’s credibility is being attacked by the media and the Democrats to preemptively head off the results of the investigation of the Clinton Foundation and Hillary Clinton. 
The covert struggle between FBI agents and Obama’s DOJ people has gone explosively public. 
The New York Times has compared Comey to J. Edgar Hoover. Its bizarre headline, “James Comey Role Recalls Hoover’s FBI, Fairly or Not” practically admits up front that it’s spouting nonsense. The Boston Globe has published a column calling for Comey’s resignation. Not to be outdone, Time has an editorial claiming that the scandal is really an attack on all women. 
James Carville appeared on MSNBC to remind everyone that he was still alive and insane. He accused Comey of coordinating with House Republicans and the KGB. And you thought the “vast right wing conspiracy” was a stretch. 
Countless media stories charge Comey with violating procedure. Do you know what’s a procedural violation? Emailing classified information stored on your bathroom server. 
Senator Harry Reid has sent Comey a letter accusing him of violating the Hatch Act. The Hatch Act is a nice idea that has as much relevance in the age of Obama as the Tenth Amendment. But the cable news spectrum quickly filled with media hacks glancing at the Wikipedia article on the Hatch Act under the table while accusing the FBI director of one of the most awkward conspiracies against Hillary ever. 
If James Comey is really out to hurt Hillary, he picked one hell of a strange way to do it. 
Not too long ago Democrats were breathing a sigh of relief when he gave Hillary Clinton a pass in a prominent public statement. If he really were out to elect Trump by keeping the email scandal going, why did he trash the investigation? Was he on the payroll of House Republicans and the KGB back then and playing it coy or was it a sudden development where Vladimir Putin and Paul Ryan talked him into taking a look at Anthony Weiner’s computer? 
Either Comey is the most cunning FBI director that ever lived or he’s just awkwardly trying to navigate a political mess that has trapped him between a DOJ leadership whose political futures are tied to Hillary’s victory and his own bureau whose apolitical agents just want to be allowed to do their jobs. 
The only truly mysterious thing is why Hillary and her associates decided to go to war with a respected Federal agency. Most Americans like the FBI while Hillary Clinton enjoys a 60% unfavorable rating. 
And it’s an interesting question. 
Hillary’s old strategy was to lie and deny that the FBI even had a criminal investigation underway. Instead her associates insisted that it was a security review. The FBI corrected her and she shrugged it off. But the old breezy denial approach has given way to a savage assault on the FBI. 
Pretending that nothing was wrong was a bad strategy, but it was a better one that picking a fight with the FBI while lunatic Clinton associates try to claim that the FBI is really the KGB. 
There are two possible explanations. 
Hillary Clinton might be arrogant enough to lash out at the FBI now that she believes that victory is near. The same kind of hubris that led her to plan her victory fireworks display could lead her to declare a war on the FBI for irritating her during the final miles of her campaign. 
But the other explanation is that her people panicked. 
Going to war with the FBI is not the behavior of a smart and focused presidential campaign. It’s an act of desperation. When a presidential candidate decides that her only option is to try and destroy the credibility of the FBI, that’s not hubris, it’s fear of what the FBI might be about to reveal about her. 
During the original FBI investigation, Hillary Clinton was confident that she could ride it out. And she had good reason for believing that. But that Hillary Clinton is gone. In her place is a paranoid wreck. Within a short space of time the “positive” Clinton campaign promising to unite the country has been replaced by a desperate and flailing operation that has focused all its energy on fighting the FBI. 
There’s only one reason for such bizarre behavior. 
The Clinton campaign has decided that an FBI investigation of the latest batch of emails poses a threat to its survival. And so it’s gone all in on fighting the FBI. It’s an unprecedented step born of fear. It’s hard to know whether that fear is justified. But the existence of that fear already tells us a whole lot. 
Clinton loyalists rigged the old investigation. They knew the outcome ahead of time as well as they knew the debate questions. Now suddenly they are no longer in control. And they are afraid. 
You can smell the fear. 
The FBI has wiretaps from the investigation of the Clinton Foundation. It’s finding new emails all the time. And Clintonworld panicked. The spinmeisters of Clintonworld have claimed that the email scandal is just so much smoke without fire. All that’s here is the appearance of impropriety without any of the substance. But this isn’t how you react to smoke. It’s how you respond to a fire. 
The misguided assault on the FBI tells us that Hillary Clinton and her allies are afraid of a revelation bigger than the fundamental illegality of her email setup. The email setup was a preemptive cover up. The Clinton campaign has panicked badly out of the belief, right or wrong, that whatever crime the illegal setup was meant to cover up is at risk of being exposed. 
The Clintons have weathered countless scandals over the years. Whatever they are protecting this time around is bigger than the usual corruption, bribery, sexual assaults and abuses of power that have followed them around throughout the years. This is bigger and more damaging than any of the allegations that have already come out. And they don’t want FBI investigators anywhere near it. 
The campaign against Comey is pure intimidation. It’s also a warning. Any senior FBI people who value their careers are being warned to stay away. The Democrats are closing ranks around their nominee against the FBI. It’s an ugly and unprecedented scene. It may also be their last stand. 
Hillary Clinton has awkwardly wound her way through numerous scandals in just this election cycle. But she’s never shown fear or desperation before. Now that has changed. Whatever she is afraid of, it lies buried in her emails with Huma Abedin. And it can bring her down like nothing else has.  
# Convert labels to numbers: REAL=1, FAKE=0
df['label_num'] = df['label'].map({'REAL': 1, 'FAKE': 0})
# Split data into training and test sets
from sklearn.model_selection import train_test_split
X = df['text']
y = df['label_num']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

What is NLP?#

NLP stands for Natural Language Processing.

It is a field of AI that helps computers understand and work with human language.

# Convert text to features with TF-IDF
from sklearn.feature_extraction.text import TfidfVectorizer
tfidf = TfidfVectorizer(stop_words='english', max_df=0.7)
X_train_tfidf = tfidf.fit_transform(X_train)
X_test_tfidf = tfidf.transform(X_test)
# Simple model: Logistic Regression
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=1000)
model.fit(X_train_tfidf, y_train)
LogisticRegression(max_iter=1000)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
# Test accuracy
score = model.score(X_test_tfidf, y_test)
print(f"Test accuracy: {score:.2f}")
Test accuracy: 0.92
# Make a fake news prediction from user input
headline = input("Type a news headline to check: ")
features = tfidf.transform([headline])
pred = model.predict(features)[0]
if pred == 1:
    print("This article looks REAL.")
else:
    print("This article looks FAKE.")
    
This article looks FAKE.

How do Transformers work?#

Transformers like BERT can understand language much more deeply than older models.

They read the whole sentence and capture the real meaning of words as they appear in context.

This is helpful in detecting sarcasm, lies, and subtle tricks in fake news.

# Use BERT transformer with Hugging Face to make predictions
from transformers import pipeline
bert_clf = pipeline('text-classification', model='distilbert-base-uncased-finetuned-sst-2-english')
result = bert_clf("Breaking news: The moon is filled with cheese!")
print(result)
[{'label': 'NEGATIVE', 'score': 0.9909403324127197}]

Practice: Try your own headline in BERT!#

Run the cell above again, but change the text to a real or fake news headline you make up.

Notice how it predicts the label.

# Visualize word importance (feature weights)
import numpy as np
import matplotlib.pyplot as plt
feature_names = np.array(tfidf.get_feature_names_out())
coefs = model.coef_[0]
top_features = 10
top_pos = np.argsort(coefs)[-top_features:]
top_neg = np.argsort(coefs)[:top_features]
plt.figure(figsize=(10, 4))
plt.barh(feature_names[top_pos], coefs[top_pos], color='green')
plt.barh(feature_names[top_neg], coefs[top_neg], color='red')
plt.xlabel('Feature weight')
plt.title('Words most linked to REAL (green) and FAKE (red) articles')
plt.show()
No description has been provided for this image

Recap#

You learned how to:

Load and explore the Fake News dataset Clean text data and encode labels Split data into training and test sets Turn headlines into features using TF-IDF Train a logistic regression model Test the model and use it to predict new headlines Try BERT transformer for deeper understanding Visualize top words for fake vs real news

Congratulations! You took your first steps towards building practical NLP tools.

Next Steps#

Try changing the classifier to Random Forest or test new headlines.

Check out other lessons for more hands-on projects!

If you learned something new, subscribe on YouTube and keep practicing!

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.