Lesson 7 · Market Research Analytics in Python
Working with Lists, Dictionaries, and Functions in Python for Market Research Analytics
This lesson will help you analyze real market research and customer survey data using Python lists, dictionaries, and functions. You will learn how to…
- CourseMarket Research Analytics in Python
- Lesson7 of 56
- Video24 min
- FormatJupyter notebook · 21 code cells
What you'll learn
Data
No separate download needed — the notebook creates or downloads everything it uses.
📓 Full notebook
Download .ipynbWorking with Lists, Dictionaries, and Functions for Customer Analytics#
- This lesson will help you analyze real market research and customer survey data using Python lists, dictionaries, and functions.
- You will learn how to manipulate customer data, group results, summarize ratings, and generate actionable insights.
- These skills are crucial for creating effective reports that drive business decisions.
- By the end, you will be able to structure and process customer datasets to answer real business questions.
import warnings
warnings.filterwarnings('ignore')
import openml
import pandas as pd
import numpy as np
About the Customer Satisfaction Survey dataset#
- This dataset contains real customer demographics and service usage information.
- Each row represents a customer's profile and their satisfaction outcome.
- Common columns include gender, senior citizen status, contract type, payment, and churn.
- Survey data must be checked for completeness and interpreted carefully to avoid incorrect conclusions.
- Beginners often mix up column meanings or forget to handle missing answers, leading to misleading insight.
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df.shape)
print(df.head(3))
# Beginner Example 1: Selecting the first 5 customer records using list slicing
customers = df.to_dict('records')
first_five = customers[:5]
print(first_five)
# Beginner Example 2: Extracting a specific column ('gender') into a list
genders = [c['gender'] for c in customers]
print(genders[:10])
# Beginner Example 3: Counting each gender using a dictionary
gender_counts = {}
for gender in genders:
if gender not in gender_counts:
gender_counts[gender] = 1
else:
gender_counts[gender] += 1
print(gender_counts)
# Beginner Example 4: Use a function to calculate the percentage of senior citizens
def senior_percentage(records):
total = len(records)
seniors = sum(1 for r in records if r['SeniorCitizen'] == 1)
return seniors / total * 100
pct = senior_percentage(customers)
print(f'Percent senior citizens: {pct:.2f}%')
# Beginner Example 5: Filter churned customers using a function and list comprehension
def filter_churned(records):
return [r for r in records if r['Churn'] == 'Yes']
churned = filter_churned(customers)
print(f'Total churned customers: {len(churned)}')
# Intermediate Example 1: Segment churned customers by contract type
def segment_by_contract(records):
segments = {}
for r in records:
ctype = r['Contract']
if ctype not in segments:
segments[ctype] = []
segments[ctype].append(r)
return segments
churned_by_contract = segment_by_contract(churned)
for k in churned_by_contract:
print(f'{k}: {len(churned_by_contract[k])} churned customers')
# Intermediate Example 2: Use a dictionary and function to calculate average tenure by gender
def average_tenure_by_gender(records):
groups = {}
for r in records:
g = r['gender']
if g not in groups:
groups[g] = []
groups[g].append(r['tenure'])
avgs = {k: np.mean(v) for k, v in groups.items()}
return avgs
result = average_tenure_by_gender(customers)
print(result)
# Intermediate Example 3: Creating a list of dictionaries for high-value customers
high_value = [r for r in customers if r['MonthlyCharges'] > 80]
print(f'Found {len(high_value)} high-value customers:')
print(high_value[:2])
# Intermediate Example 4: Define a reusable function to summarize categorical distributions
def summarize_category(records, key):
counts = {}
for r in records:
v = r[key]
if v not in counts:
counts[v] = 1
else:
counts[v] += 1
total = sum(counts.values())
summary = {k: f'{v} ({v/total*100:.1f}%)' for k, v in counts.items()}
return summary
contract_summary = summarize_category(customers, 'Contract')
print(contract_summary)
# Intermediate Example 5: Map 'Yes'/'No' churn to 1/0 using dictionary comprehension
churn_map = {'Yes': 1, 'No': 0}
churned_numeric = [churn_map.get(r['Churn'], np.nan) for r in customers]
print(churned_numeric[:10])
# Advanced Example 1: Cross-tabulate gender and churn using a nested dictionary
cross_tab = {}
for r in customers:
g = r['gender']
c = r['Churn']
if g not in cross_tab:
cross_tab[g] = {}
if c not in cross_tab[g]:
cross_tab[g][c] = 0
cross_tab[g][c] += 1
print(cross_tab)
# Advanced Example 2: Build a function that groups customers by a field and calculates a custom metric
def aggregate_metric(records, group_field, metric_func):
groups = {}
for r in records:
key = r[group_field]
if key not in groups:
groups[key] = []
groups[key].append(r)
summary = {k: metric_func(v) for k, v in groups.items()}
return summary
def avg_monthly_charge(records):
return np.mean([r['MonthlyCharges'] for r in records])
charge_summary = aggregate_metric(customers, 'Contract', avg_monthly_charge)
print(charge_summary)
# Advanced Example 3: Multi-step pipeline using lists, dictionaries, and functions to profile churn risk
def high_risk_profile(records, charge_thresh=70, max_tenure=12):
return [r for r in records if r['MonthlyCharges'] > charge_thresh and r['tenure'] < max_tenure and r['Churn']=='Yes']
risk_customers = high_risk_profile(customers)
risk_segments = summarize_category(risk_customers, 'Contract')
print('High-churn, high-value segments:', risk_segments)
# Advanced Example 4: Write and read a summary report file
with open('customer_churn_summary.txt', 'w') as f:
for k,v in charge_summary.items():
f.write(f'{k}: Average Monthly Charge ${v:.2f}\n')
for contract, v in risk_segments.items():
f.write(f'High Risk - {contract}: {v}\n')
# Error Handling Example 1: Handle missing survey responses ('TotalCharges' as blank)
num_blank = sum(1 for r in customers if r.get('TotalCharges', '') == '')
print(f'Total customers missing TotalCharges: {num_blank}')
# Error Handling Example 2: Prevent crash on incorrect group field by using dict.get
k = 'InvalidField'
try:
value = customers[0][k]
except KeyError:
value = None
print('Sample value:', value)
# Error Handling Example 3: Detect and warn about misinterpreted NPS or Likert scales
def check_likert(nps_list):
unexpected = [v for v in nps_list if v < 0 or v > 10]
if unexpected:
print(f'Warning: {len(unexpected)} NPS values are outside of 0-10 scale')
else:
print('All NPS scores within valid range')
np.random.seed(42)
nps_scores = list(np.random.randint(-2,13,25))
check_likert(nps_scores)
Best Practices in Market Research with Lists, Dictionaries, and Functions#
- Use clear segmentation functions to compare trends across groups.
- For cross-tabulations, nest dictionaries for each factor (e.g., contract type by churn).
- When making a trend index, combine counts and percents to explain both magnitude and risk.
- For scoring, convert text responses or Yes/No into numerical indicators early.
- Trend analysis often needs results stored in a list (for ordering) and summary dictionaries (for business reporting).
# End-to-end Market Research Problem: Predicting churn risk for monthly contract customers
monthly_customers = [r for r in customers if r['Contract'] == 'Month-to-month']
risk_metric = lambda r: int(r['MonthlyCharges'] > 75 and r['tenure'] < 6)
risks = [risk_metric(r) for r in monthly_customers]
risk_rate = sum(risks) / len(monthly_customers) * 100 if monthly_customers else 0
print(f'Monthly contract, high-risk churn rate: {risk_rate:.1f}%')
Found this useful?
All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.



