Lesson 10 · OpenAI
How to Use OpenAI Whisper for Accurate Speech-to-Text Transcription
Welcome to your hands-on beginner lesson. We will explore the OpenAI Python API together. You will learn how to use the API for tasks like chat, speech to…
- CourseOpenAI
- Lesson10 of 14
- Video18 min
- FormatJupyter notebook · 14 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbIntroduction to the OpenAI API in Python#
- Welcome to your hands-on beginner lesson.
- We will explore the OpenAI Python API together.
- You will learn how to use the API for tasks like chat, speech to text, and more.
- Let us get started!
What is the OpenAI API?#
- The OpenAI API lets you use advanced AI models in your own programs.
- You can generate text, transcribe speech, or create embeddings.
- Many popular apps use this API to add smart features.
- You do not need deep AI knowledge to begin.
Why Learn the OpenAI API?#
- AI can save time and open creative possibilities.
- Automate tasks like summarizing text or converting speech to text.
- Improve user experiences in apps with natural conversations.
- The skills you learn here will help in many fields.
Getting Set Up#
- You need a Python environment and an OpenAI API key.
- This lesson runs in Jupyter, which is great for experiments.
- The API key is already in the environment for you as OPENAI_API_KEY.
import warnings; warnings.filterwarnings("ignore") # This hides warning messages.
import numpy as np
np.random.seed(42)
import os # Allows us to use operating system features.
from openai import OpenAI # This is the main OpenAI client library.
API_key = os.environ["OPENAI_API_KEY"] # Fetch our API key from the environment.
client = OpenAI(api_key=API_key) # Create the OpenAI client with your key.
Key Concepts#
- Models are different AI brains for different jobs.
- Prompts are messages you send that the AI responds to.
- Always handle API keys with care.
- Cost depends on the model and usage.
- Never share your key publicly.
# Let us try a basic Chat Completion.
response = client.chat.completions.create(
model="gpt-4.1-mini",
messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content) # Print the response text.
How Chat Completions Work#
- Each message has a role like "user" or "assistant".
- The model reads your messages and replies in context.
- You can choose a faster, cheaper model or a stronger model depending on your needs.
- You can create smart chatbots or assistants.
# Use input() to send your own chat prompt.
user_msg = input("Type your message for the assistant: ")
my_response = client.chat.completions.create(
model="gpt-4.1-mini",
messages=[{"role": "user", "content": user_msg}]
)
print(my_response.choices[0].message.content)
Speech to Text with Whisper#
- You can turn spoken words into written text using OpenAI Whisper.
- This is called speech recognition.
- It is great for notes, captions, and accessibility.
# Prepare to use speech recognition.
audio_file = open("sample.wav", "rb") # Open an audio file. Replace with your file if needed.
transcript = client.audio.transcriptions.create(
model="whisper-1",
file=audio_file
)
print(transcript.text) # Show the transcribed text.
# Try transcribing a different audio file with input().
file_path = input("Enter the path to your audio file: ")
with open(file_path, "rb") as my_audio:
my_transcription = client.audio.transcriptions.create(
model="whisper-1",
file=my_audio
)
print(my_transcription.text)
Embeddings: Making Text Searchable#
- Embeddings turn text into lists of numbers called vectors.
- You can use these vectors to compare meaning, power search, or cluster ideas.
- Embeddings help AI organize language by meaning rather than spelling.
# Create an embedding for a simple sentence.
embed = client.embeddings.create(
model="text-embedding-3-small",
input=["Machine learning is fun"]
)
print(len(embed.data[0].embedding)) # Shows the vector size.
# Compare two sentences using their embeddings.
import numpy as np # Lets us do math with arrays.
sentences = ["OpenAI makes AI tools.", "Artificial intelligence tools by OpenAI."]
embs = client.embeddings.create(
model="text-embedding-3-small",
input=sentences
)
a, b = np.array(embs.data[0].embedding), np.array(embs.data[1].embedding)
# Use cosine similarity to compare.
similarity = np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
print("Similarity:", similarity)
Error Handling#
- Always use try-except blocks when calling the API.
- This helps prevent crashes if there is a network error or a typo.
- Good error handling is important for real-world apps.
# Wrap a chat completion call with error handling.
try:
resp = client.chat.completions.create(
model="gpt-4.1-mini",
messages=[{"role": "user", "content": "Give me a joke."}]
)
print(resp.choices[0].message.content)
except Exception as e:
print("Error: ", e)
Best Practices#
- Use the smallest model that does the job to save money.
- Limit how often you make API calls.
- Never share API keys or put them on GitHub.
- Check model and API documentation for updates.
# Mini Project: Turn an audio file to text and summarize it using chat.
audio_file = open("sample.wav", "rb")
transcript = client.audio.transcriptions.create(
model="whisper-1",
file=audio_file
)
summary = client.chat.completions.create(
model="gpt-4.1-mini",
messages=[
{"role": "system", "content": "You summarize transcripts."},
{"role": "user", "content": transcript.text}
]
)
print("Summary:\n", summary.choices[0].message.content)
Troubleshooting Tips#
- If you see 'API key error', check your environment variable.
- For connection errors, check your internet.
- If you get usage or rate limit errors, try again later or use a smaller model.
- Read error messages carefullythey will tell you what to fix.
# Challenge 1: Try summarizing your own audio file.
file_path = input("Type the path to your own audio file: ")
with open(file_path, "rb") as my_audio:
my_transcript = client.audio.transcriptions.create(
model="whisper-1",
file=my_audio
)
my_summary = client.chat.completions.create(
model="gpt-4.1-mini",
messages=[
{"role": "system", "content": "You summarize transcripts."},
{"role": "user", "content": my_transcript.text}
]
)
print("Summary:\n", my_summary.choices[0].message.content)
# Challenge 2: Get keyword ideas using chat completions.
topic = input("Enter a topic or theme: ")
ideas = client.chat.completions.create(
model="gpt-4.1-mini",
messages=[
{"role": "system", "content": "You produce a list of keywords."},
{"role": "user", "content": topic}
]
)
print("Keywords:\n", ideas.choices[0].message.content)
Recap#
- You learned how to set up and use the OpenAI API in Python.
- We covered chat completions, speech to text, and embeddings.
- You tried mini projects and challenge exercises.
- Now you are ready to build your own smart apps!
Thanks and Next Steps#
- Congratulations! Now you have the basics of the OpenAI API in Python.
- Try using these features in your homework or hobby projects.
- Check the OpenAI docs for advanced examples.
- Like and subscribe for more beginner AI lessons!
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



