Lesson 7 · OpenAI
Understanding Vector Search and Semantic Similarity in AI and Machine Learning
Welcome to your beginner friendly guide to using the OpenAI API. This lesson shows you how you can talk to advanced AI using code. You will learn the…
- CourseOpenAI
- Lesson7 of 14
- Video25 min
- FormatJupyter notebook · 16 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbGetting Started with the OpenAI API in Python#
- Welcome to your beginner friendly guide to using the OpenAI API.
- This lesson shows you how you can talk to advanced AI using code.
- You will learn the basics, key safety tips, and create simple AI helpers.
- By the end, you can use Python to access, chat, and explore AI features.
What is the OpenAI API?#
- The OpenAI API is a powerful tool for working with advanced artificial intelligence.
- It lets your Python code generate text, answer questions, summarize, and even understand audio and images.
- The API connects your computer to AI models like ChatGPT using simple requests.
- Anyone can use it to build intelligent apps, creative tools, and much more.
Setting up Python and the OpenAI Client Library#
- Before you start, you must install the official OpenAI package for Python.
- This package lets your code easily talk to the OpenAI API.
- You can use pip, a tool for installing Python libraries, to do this.
import numpy as np
np.random.seed(42)
# Install the OpenAI package if needed
get_ipython().system('pip install --upgrade openai')
# Suppress warning messages for a cleaner experience
import warnings; warnings.filterwarnings("ignore")
# Import the OpenAI client class
from openai import OpenAI
# Import os for reading our API key from the environment
import os
Staying Secure with API Keys#
- OpenAI requires an API key, which is like a password for your account.
- You must keep this key safe and never share it inside your code.
- For these lessons, we already saved the API key securely in the environment.
- You do not need to enter it or paste it anywhere.
- Your code will load it automatically each time.
# Read your OpenAI API key from the environment (never put it in your code!)
API_key = os.environ["OPENAI_API_KEY"]
# Create a client instance to connect with the OpenAI API
client = OpenAI()
Understanding Chat Completions#
- The most popular way to use the OpenAI API is through chat completions.
- In simple words, you send messages to an AI and get helpful answers.
- You choose the model, write a question, and receive a smart reply.
- This is the same core power behind ChatGPT.
# Send a simple message and get a reply from the AI
response = client.chat.completions.create(
model='gpt-4.1-mini',
messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)
# Send your own question to the AI
user_question = input('Enter any question for AI: ')
response = client.chat.completions.create(
model='gpt-4.1-mini',
messages=[{"role": "user", "content": user_question}]
)
print('AI Response:', response.choices[0].message.content)
# What is semantic similarity?
explanation = client.chat.completions.create(
model='gpt-4.1-mini',
messages=[{"role": "user", "content": "Explain semantic similarity in simple words"}]
)
print(explanation.choices[0].message.content)
What Are Embeddings and Vector Similarity?#
- Embeddings turn your text into a list of numbers that the AI uses to compare ideas.
- These numbers are called vectors, and they help the AI understand if texts are similar.
- You can use embeddings to match search results, organize data, or cluster meanings.
- Vector similarity means comparing how close two pieces of text are, using math.
# Let us create embeddings for some text
embed = client.embeddings.create(
model='text-embedding-3-small',
input=['Machine learning is fun']
)
embedding_vector = embed.data[0].embedding
print('Vector length:', len(embedding_vector))
# Compare the semantic similarity of two short sentences
def cosine_similarity(vec1, vec2):
"""Compute cosine similarity between two lists of numbers."""
import math
dot = sum(a*b for a,b in zip(vec1, vec2))
norm1 = math.sqrt(sum(a*a for a in vec1))
norm2 = math.sqrt(sum(b*b for b in vec2))
if norm1 and norm2:
return dot / (norm1 * norm2)
return 0
text1 = 'Cats love to nap in the sun'
text2 = 'Felines like sleeping where it is warm'
# Get embeddings for both texts
res = client.embeddings.create(
model='text-embedding-3-small',
input=[text1, text2]
)
vec1 = res.data[0].embedding
vec2 = res.data[1].embedding
# Show similarity score
score = cosine_similarity(vec1, vec2)
print('Similarity score:', round(score, 2))
How Can You Use Similarity in Real Life?#
- Semantic similarity helps you find duplicate questions, cluster users by their interests, or even build smarter search engines.
- Vector-based tools are more flexible than exact keyword matching.
- You can power up chatbots, study documents, or analyze feedback meaningfully.
# Mini project: Find the most similar FAQ for a user question
faqs = [
'How do I reset my password?',
'What are your working hours?',
'Where can I update my email preferences?',
'How can I contact support?'
]
user_query = input('Ask a help question: ')
# Combine user question with FAQs for batch embedding
all_texts = [user_query] + faqs
emb = client.embeddings.create(
model='text-embedding-3-small',
input=all_texts
)
user_vec = emb.data[0].embedding
faq_vecs = [item.embedding for item in emb.data[1:]]
# Calculate similarity with each FAQ
scores = [cosine_similarity(user_vec, v) for v in faq_vecs]
best_match = faqs[scores.index(max(scores))]
print('Best match:', best_match)
# What if there is no strong match?
user_question2 = input('Enter another question: ')
all_texts2 = [user_question2] + faqs
emb2 = client.embeddings.create(
model='text-embedding-3-small',
input=all_texts2
)
user_vec2 = emb2.data[0].embedding
faq_vecs2 = [item.embedding for item in emb2.data[1:]]
scores2 = [cosine_similarity(user_vec2, v) for v in faq_vecs2]
best_score = max(scores2)
if best_score > 0.5:
print('Best match:', faqs[scores2.index(best_score)])
else:
print('Sorry, no close match found.')
# Error handling: What if you hit your API rate limit?
try:
broken = client.chat.completions.create(
model='gpt-4.1-mini',
messages=[{"role": "user", "content": "Test rate limit"}]
)
print(broken.choices[0].message.content)
except Exception as e:
print('There was an error:', str(e))
# Best practices: Use smaller models when possible
small_model = 'gpt-4.1-mini'
large_model = 'gpt-4.1'
print('Choose small models for quick tests. Use larger models for more complex questions.')
# Challenge: Try asking for three similar phrases and check their similarities
sentence1 = input('First phrase: ')
sentence2 = input('Second phrase: ')
sentence3 = input('Third phrase: ')
input_list = [sentence1, sentence2, sentence3]
embeds = client.embeddings.create(
model='text-embedding-3-small',
input=input_list
)
vecs = [d.embedding for d in embeds.data]
for i in range(3):
for j in range(i+1, 3):
sim = cosine_similarity(vecs[i], vecs[j])
print(f"Similarity between {i+1} and {j+1}: {round(sim,2)}")
# Challenge: Try a phrase and see if it matches a completely unrelated FAQ
unrel_text = input('Enter a random topic: ')
to_check = [unrel_text] + faqs
unrel_emb = client.embeddings.create(
model='text-embedding-3-small',
input=to_check
)
main_vec = unrel_emb.data[0].embedding
comp_vecs = [d.embedding for d in unrel_emb.data[1:]]
sim_scores = [cosine_similarity(main_vec, v) for v in comp_vecs]
print('Most similar FAQ:', faqs[sim_scores.index(max(sim_scores))])
print('Highest similarity score:', round(max(sim_scores), 2))
Recap: You Have Learned the Basics#
- You can now connect to the OpenAI API, send chat prompts, and get generated answers.
- You explored embeddings and vector similarity to compare how much meaning two texts share.
- By combining these skills, you can build creative and helpful AI applications.
Keep Exploring and Subscribe for More!#
- Try changing the model and repeating the challenge exercises above.
- Experiment with different questions or build your own FAQ lists.
- Practice makes you more comfortable with AI and code.
- Subscribe for more beginner tutorials and projects.
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



