Python library centre
How to Train and Fine-Tune NLP Models Using SpaCy: A Step-by-Step Guide
SpaCy is a powerful Python library for Natural Language Processing (NLP). It is used to work with human language data easily. Common real-world uses are…
- CoursePython library centre
- Video19 min
- FormatJupyter notebook · 21 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIntroduction to SpaCy#
SpaCy is a powerful Python library for Natural Language Processing (NLP).
It is used to work with human language data easily.
Common real-world uses are automatic text classification, extracting information, and analyzing documents.
SpaCy makes it easy to process text, split text into words, and understand language structure.
In this lesson, we will learn the basics of SpaCy step-by-step.
import warnings; warnings.filterwarnings("ignore")
# To install: open your command prompt and run:
# pip install spacy
# You will also need to download a model:
# python -m spacy download en_core_web_sm
import spacy
from spacy.language import Language
Core Objects in SpaCy#
- The main object is the 'nlp' object. It processes your text.
- When you pass text to the nlp object, you get a 'Doc' object.
- A Doc object holds your text, words, and linguistic information.
- Tokens represent words or punctuation.
- The library provides access to Parts of Speech, Lemmas, Entities, and more.
nlp = spacy.load("en_core_web_sm")
text = "SpaCy is an awesome library for NLP!"
doc = nlp(text)
print(doc)
for token in doc:
print(token.text)
for token in doc:
print(f"{token.text}: {token.pos_}")
Example: Tokenization#
- Tokenization means splitting text into words and punctuation.
- SpaCy handles tokenization automatically.
- Each token is an object with helpful properties.
sentence = "Python is great for data science!"
doc2 = nlp(sentence)
for token in doc2:
print(token.text, token.idx)
print("Has vector:", doc2.has_vector)
print("Is tagged:", doc2.is_tagged)
print("Is parsed:", doc2.is_parsed)
Example: Lemmatization#
- Lemmatization reduces words to their base forms.
- For example, 'running' becomes 'run'.
- SpaCy makes it easy to get the lemma for each word.
lemmas_example = "I am running and I ran every day."
doc3 = nlp(lemmas_example)
for token in doc3:
print(f"{token.text} -> {token.lemma_}")
Example: Named Entity Recognition (NER)#
- NER finds names, places, dates, and organizations in text.
- SpaCy can identify named entities easily.
ner_sentence = "Apple was founded by Steve Jobs in California in 1976."
doc4 = nlp(ner_sentence)
for ent in doc4.ents:
print(ent.text, ent.label_)
Example: Stop Words#
- Stop words are common words like 'the', 'is', 'and'.
- Often, these words are ignored in text processing.
- SpaCy provides a list of stop words.
for token in doc4:
if token.is_stop:
print(token.text)
Intermediate: Dependency Parsing#
- SpaCy understands how words are linked (syntax).
- You can see the relation between words using dependencies.
sent = "The quick brown fox jumps over the lazy dog."
doc5 = nlp(sent)
for token in doc5:
print(token.text, "<-", token.dep_, "-", token.head.text)
for chunk in doc5.noun_chunks:
print(chunk.text)
Intermediate: Sentence Segmentation#
- SpaCy splits text into sentences automatically.
- These sentences can be processed separately.
multi_sent = "SpaCy is easy. It is fast. It is accurate."
doc6 = nlp(multi_sent)
for sent in doc6.sents:
print(sent.text)
Intermediate: Custom Stop Words#
- You can add your own stop words to SpaCy.
- This can help handle domain-specific needs.
nlp.Defaults.stop_words.add("awesome")
doc7 = nlp("SpaCy is awesome and useful.")
for token in doc7:
if token.is_stop:
print(token.text)
Advanced: Create Custom Pipeline Component#
- You can add your own processing step into SpaCy's NLP pipeline.
- This allows you to run custom functions on all texts.
@Language.component("custom_component")
def custom_component(doc):
print("Custom component ran!")
return doc
nlp.add_pipe("custom_component", last=True)
doc8 = nlp("Testing component.")
Advanced: Save and Load SpaCy Models#
- You can save your trained model or modified pipeline.
- Later, you can load it back for use.
nlp.to_disk("my_spacy_pipeline") # Save pipeline to folder
loaded_nlp = spacy.load("my_spacy_pipeline")
doc9 = loaded_nlp("Model loaded!")
print(doc9.text)
Error Handling in SpaCy#
- SpaCy will raise errors if you use a model that is not downloaded.
- Let us handle some typical errors with try / except.
try:
faulty_nlp = spacy.load("nonexistent_model")
except Exception as e:
print("Error:", e)
try:
blank_nlp = spacy.blank("zz")
except Exception as e:
print("Error initializing blank model:", e)
Best Practices with SpaCy#
- Always use the smallest model possible for faster results.
- Use pip requirements.txt to manage packages.
- Save and load pipelines to reuse your setup.
- Handle exceptions to make programs robust.
- Read the SpaCy documentation for updates and new features.
def is_entity(text):
doc = nlp(text)
if len(doc.ents) > 0:
print("Entities detected:")
for ent in doc.ents:
print(ent.text, ent.label_)
else:
print("No entities found.")
is_entity("Barack Obama was President of the United States.")
is_entity("Nothing special here.")
Mini Project: Extract Emails from Text#
- Let us build a function to extract email addresses using SpaCy.
- This is useful for contact scraping and document analysis.
import re
def extract_emails(text):
pattern = r"[\w.-]+@[\w.-]+\.[a-zA-Z]{2,}"
emails = re.findall(pattern, text)
if emails:
print("Emails found:", emails)
else:
print("No emails found.")
sample = "Contact us at info@example.com or support@mail.com."
extract_emails(sample)
YouTube Call to Action#
- Please subscribe to our channel for more tutorials!
- Like this video if you found it helpful.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



