Mathew K Analytics

Python library centre

How to Train and Fine-Tune NLP Models Using SpaCy: A Step-by-Step Guide

SpaCy is a powerful Python library for Natural Language Processing (NLP). It is used to work with human language data easily. Common real-world uses are…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Introduction to SpaCy#

  • SpaCy is a powerful Python library for Natural Language Processing (NLP).

  • It is used to work with human language data easily.

  • Common real-world uses are automatic text classification, extracting information, and analyzing documents.

  • SpaCy makes it easy to process text, split text into words, and understand language structure.

  • In this lesson, we will learn the basics of SpaCy step-by-step.

import warnings; warnings.filterwarnings("ignore")
 
# To install: open your command prompt and run:
# pip install spacy
 
# You will also need to download a model:
# python -m spacy download en_core_web_sm
 
import spacy
from spacy.language import Language

Core Objects in SpaCy#

  • The main object is the 'nlp' object. It processes your text.
  • When you pass text to the nlp object, you get a 'Doc' object.
  • A Doc object holds your text, words, and linguistic information.
  • Tokens represent words or punctuation.
  • The library provides access to Parts of Speech, Lemmas, Entities, and more.
nlp = spacy.load("en_core_web_sm")
 
text = "SpaCy is an awesome library for NLP!"
doc = nlp(text)
 
print(doc)
SpaCy is an awesome library for NLP!
for token in doc:
    print(token.text)
SpaCy
is
an
awesome
library
for
NLP
!
for token in doc:
    print(f"{token.text}: {token.pos_}")
SpaCy: PROPN
is: AUX
an: DET
awesome: ADJ
library: NOUN
for: ADP
NLP: PROPN
!: PUNCT

Example: Tokenization#

  • Tokenization means splitting text into words and punctuation.
  • SpaCy handles tokenization automatically.
  • Each token is an object with helpful properties.
sentence = "Python is great for data science!"
doc2 = nlp(sentence)
 
for token in doc2:
    print(token.text, token.idx)
Python 0
is 7
great 10
for 16
data 20
science 25
! 32
print("Has vector:", doc2.has_vector)
print("Is tagged:", doc2.is_tagged)
print("Is parsed:", doc2.is_parsed)
Has vector: True
Is tagged: True
Is parsed: True

Example: Lemmatization#

  • Lemmatization reduces words to their base forms.
  • For example, 'running' becomes 'run'.
  • SpaCy makes it easy to get the lemma for each word.
lemmas_example = "I am running and I ran every day."
doc3 = nlp(lemmas_example)
 
for token in doc3:
    print(f"{token.text} -> {token.lemma_}")
I -> I
am -> be
running -> run
and -> and
I -> I
ran -> run
every -> every
day -> day
. -> .

Example: Named Entity Recognition (NER)#

  • NER finds names, places, dates, and organizations in text.
  • SpaCy can identify named entities easily.
ner_sentence = "Apple was founded by Steve Jobs in California in 1976."
doc4 = nlp(ner_sentence)
for ent in doc4.ents:
    print(ent.text, ent.label_)
Apple ORG
Steve Jobs PERSON
California GPE
1976 DATE

Example: Stop Words#

  • Stop words are common words like 'the', 'is', 'and'.
  • Often, these words are ignored in text processing.
  • SpaCy provides a list of stop words.
for token in doc4:
    if token.is_stop:
        print(token.text)
was
by
in
in

Intermediate: Dependency Parsing#

  • SpaCy understands how words are linked (syntax).
  • You can see the relation between words using dependencies.
sent = "The quick brown fox jumps over the lazy dog."
doc5 = nlp(sent)
for token in doc5:
    print(token.text, "<-", token.dep_, "-", token.head.text)
The <- det - fox
quick <- amod - fox
brown <- amod - fox
fox <- nsubj - jumps
jumps <- ROOT - jumps
over <- prep - jumps
the <- det - dog
lazy <- amod - dog
dog <- pobj - over
. <- punct - jumps
for chunk in doc5.noun_chunks:
    print(chunk.text)
The quick brown fox
the lazy dog

Intermediate: Sentence Segmentation#

  • SpaCy splits text into sentences automatically.
  • These sentences can be processed separately.
multi_sent = "SpaCy is easy. It is fast. It is accurate."
doc6 = nlp(multi_sent)
for sent in doc6.sents:
    print(sent.text)
SpaCy is easy.
It is fast.
It is accurate.

Intermediate: Custom Stop Words#

  • You can add your own stop words to SpaCy.
  • This can help handle domain-specific needs.
nlp.Defaults.stop_words.add("awesome")
doc7 = nlp("SpaCy is awesome and useful.")
for token in doc7:
    if token.is_stop:
        print(token.text)
is
and

Advanced: Create Custom Pipeline Component#

  • You can add your own processing step into SpaCy's NLP pipeline.
  • This allows you to run custom functions on all texts.
@Language.component("custom_component")
def custom_component(doc):
    print("Custom component ran!")
    return doc
 
nlp.add_pipe("custom_component", last=True)
doc8 = nlp("Testing component.")
Custom component ran!

Advanced: Save and Load SpaCy Models#

  • You can save your trained model or modified pipeline.
  • Later, you can load it back for use.
nlp.to_disk("my_spacy_pipeline")  # Save pipeline to folder
loaded_nlp = spacy.load("my_spacy_pipeline")
doc9 = loaded_nlp("Model loaded!")
print(doc9.text)
Custom component ran!
Model loaded!

Error Handling in SpaCy#

  • SpaCy will raise errors if you use a model that is not downloaded.
  • Let us handle some typical errors with try / except.
try:
    faulty_nlp = spacy.load("nonexistent_model")
except Exception as e:
    print("Error:", e)
Error: [E050] Can't find model 'nonexistent_model'. It doesn't seem to be a Python package or a valid path to a data directory.
try:
    blank_nlp = spacy.blank("zz")
except Exception as e:
    print("Error initializing blank model:", e)
Error initializing blank model: [E048] Can't import language zz or any matching language from spacy.lang: No module named 'spacy.lang.zz'

Best Practices with SpaCy#

  • Always use the smallest model possible for faster results.
  • Use pip requirements.txt to manage packages.
  • Save and load pipelines to reuse your setup.
  • Handle exceptions to make programs robust.
  • Read the SpaCy documentation for updates and new features.
def is_entity(text):
    doc = nlp(text)
    if len(doc.ents) > 0:
        print("Entities detected:")
        for ent in doc.ents:
            print(ent.text, ent.label_)
    else:
        print("No entities found.")
is_entity("Barack Obama was President of the United States.")
is_entity("Nothing special here.")
Custom component ran!
Entities detected:
Barack Obama PERSON
the United States GPE
Custom component ran!
No entities found.

Mini Project: Extract Emails from Text#

  • Let us build a function to extract email addresses using SpaCy.
  • This is useful for contact scraping and document analysis.
import re
def extract_emails(text):
    pattern = r"[\w.-]+@[\w.-]+\.[a-zA-Z]{2,}"
    emails = re.findall(pattern, text)
    if emails:
        print("Emails found:", emails)
    else:
        print("No emails found.")
sample = "Contact us at info@example.com or support@mail.com."
extract_emails(sample)
Emails found: ['info@example.com', 'support@mail.com']

YouTube Call to Action#

  • Please subscribe to our channel for more tutorials!
  • Like this video if you found it helpful.

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.