Mathew K Analytics

Lesson 7 · Market Research Analytics in Python

Working with Lists, Dictionaries, and Functions in Python for Market Research Analytics

This lesson will help you analyze real market research and customer survey data using Python lists, dictionaries, and functions. You will learn how to…

⬇ Download notebookOpen in Colab ↗

What you'll learn

Data

No separate download needed — the notebook creates or downloads everything it uses.

📓 Full notebook

Download .ipynb

Working with Lists, Dictionaries, and Functions for Customer Analytics#

  • This lesson will help you analyze real market research and customer survey data using Python lists, dictionaries, and functions.
  • You will learn how to manipulate customer data, group results, summarize ratings, and generate actionable insights.
  • These skills are crucial for creating effective reports that drive business decisions.
  • By the end, you will be able to structure and process customer datasets to answer real business questions.
import warnings
warnings.filterwarnings('ignore')
import openml
import pandas as pd
import numpy as np

About the Customer Satisfaction Survey dataset#

  • This dataset contains real customer demographics and service usage information.
  • Each row represents a customer's profile and their satisfaction outcome.
  • Common columns include gender, senior citizen status, contract type, payment, and churn.
  • Survey data must be checked for completeness and interpreted carefully to avoid incorrect conclusions.
  • Beginners often mix up column meanings or forget to handle missing answers, leading to misleading insight.
dataset = openml.datasets.get_dataset(42178)
df, _, _, _ = dataset.get_data(dataset_format='dataframe')
print(df.shape)
print(df.head(3))
(7043, 20)
   gender  SeniorCitizen Partner Dependents  tenure PhoneService  \
0  Female              0     Yes         No       1           No   
1    Male              0      No         No      34          Yes   
2    Male              0      No         No       2          Yes   

      MultipleLines InternetService OnlineSecurity OnlineBackup  \
0  No phone service             DSL             No          Yes   
1                No             DSL            Yes           No   
2                No             DSL            Yes          Yes   

  DeviceProtection TechSupport StreamingTV StreamingMovies        Contract  \
0               No          No          No              No  Month-to-month   
1              Yes          No          No              No        One year   
2               No          No          No              No  Month-to-month   

  PaperlessBilling     PaymentMethod  MonthlyCharges TotalCharges Churn  
0              Yes  Electronic check           29.85        29.85    No  
1               No      Mailed check           56.95       1889.5    No  
2              Yes      Mailed check           53.85       108.15   Yes  
# Beginner Example 1: Selecting the first 5 customer records using list slicing
customers = df.to_dict('records')
first_five = customers[:5]
print(first_five)
[{'gender': 'Female', 'SeniorCitizen': 0, 'Partner': 'Yes', 'Dependents': 'No', 'tenure': 1, 'PhoneService': 'No', 'MultipleLines': 'No phone service', 'InternetService': 'DSL', 'OnlineSecurity': 'No', 'OnlineBackup': 'Yes', 'DeviceProtection': 'No', 'TechSupport': 'No', 'StreamingTV': 'No', 'StreamingMovies': 'No', 'Contract': 'Month-to-month', 'PaperlessBilling': 'Yes', 'PaymentMethod': 'Electronic check', 'MonthlyCharges': 29.85, 'TotalCharges': '29.85', 'Churn': 'No'}, {'gender': 'Male', 'SeniorCitizen': 0, 'Partner': 'No', 'Dependents': 'No', 'tenure': 34, 'PhoneService': 'Yes', 'MultipleLines': 'No', 'InternetService': 'DSL', 'OnlineSecurity': 'Yes', 'OnlineBackup': 'No', 'DeviceProtection': 'Yes', 'TechSupport': 'No', 'StreamingTV': 'No', 'StreamingMovies': 'No', 'Contract': 'One year', 'PaperlessBilling': 'No', 'PaymentMethod': 'Mailed check', 'MonthlyCharges': 56.95, 'TotalCharges': '1889.5', 'Churn': 'No'}, {'gender': 'Male', 'SeniorCitizen': 0, 'Partner': 'No', 'Dependents': 'No', 'tenure': 2, 'PhoneService': 'Yes', 'MultipleLines': 'No', 'InternetService': 'DSL', 'OnlineSecurity': 'Yes', 'OnlineBackup': 'Yes', 'DeviceProtection': 'No', 'TechSupport': 'No', 'StreamingTV': 'No', 'StreamingMovies': 'No', 'Contract': 'Month-to-month', 'PaperlessBilling': 'Yes', 'PaymentMethod': 'Mailed check', 'MonthlyCharges': 53.85, 'TotalCharges': '108.15', 'Churn': 'Yes'}, {'gender': 'Male', 'SeniorCitizen': 0, 'Partner': 'No', 'Dependents': 'No', 'tenure': 45, 'PhoneService': 'No', 'MultipleLines': 'No phone service', 'InternetService': 'DSL', 'OnlineSecurity': 'Yes', 'OnlineBackup': 'No', 'DeviceProtection': 'Yes', 'TechSupport': 'Yes', 'StreamingTV': 'No', 'StreamingMovies': 'No', 'Contract': 'One year', 'PaperlessBilling': 'No', 'PaymentMethod': 'Bank transfer (automatic)', 'MonthlyCharges': 42.3, 'TotalCharges': '1840.75', 'Churn': 'No'}, {'gender': 'Female', 'SeniorCitizen': 0, 'Partner': 'No', 'Dependents': 'No', 'tenure': 2, 'PhoneService': 'Yes', 'MultipleLines': 'No', 'InternetService': 'Fiber optic', 'OnlineSecurity': 'No', 'OnlineBackup': 'No', 'DeviceProtection': 'No', 'TechSupport': 'No', 'StreamingTV': 'No', 'StreamingMovies': 'No', 'Contract': 'Month-to-month', 'PaperlessBilling': 'Yes', 'PaymentMethod': 'Electronic check', 'MonthlyCharges': 70.7, 'TotalCharges': '151.65', 'Churn': 'Yes'}]
# Beginner Example 2: Extracting a specific column ('gender') into a list
genders = [c['gender'] for c in customers]
print(genders[:10])
['Female', 'Male', 'Male', 'Male', 'Female', 'Female', 'Male', 'Female', 'Female', 'Male']
# Beginner Example 3: Counting each gender using a dictionary
gender_counts = {}
for gender in genders:
    if gender not in gender_counts:
        gender_counts[gender] = 1
    else:
        gender_counts[gender] += 1
print(gender_counts)
{'Female': 3488, 'Male': 3555}
# Beginner Example 4: Use a function to calculate the percentage of senior citizens
def senior_percentage(records):
    total = len(records)
    seniors = sum(1 for r in records if r['SeniorCitizen'] == 1)
    return seniors / total * 100

pct = senior_percentage(customers)
print(f'Percent senior citizens: {pct:.2f}%')
Percent senior citizens: 16.21%
# Beginner Example 5: Filter churned customers using a function and list comprehension
def filter_churned(records):
    return [r for r in records if r['Churn'] == 'Yes']

churned = filter_churned(customers)
print(f'Total churned customers: {len(churned)}')
Total churned customers: 1869
# Intermediate Example 1: Segment churned customers by contract type
def segment_by_contract(records):
    segments = {}
    for r in records:
        ctype = r['Contract']
        if ctype not in segments:
            segments[ctype] = []
        segments[ctype].append(r)
    return segments

churned_by_contract = segment_by_contract(churned)
for k in churned_by_contract:
    print(f'{k}: {len(churned_by_contract[k])} churned customers')
Month-to-month: 1655 churned customers
Two year: 48 churned customers
One year: 166 churned customers
# Intermediate Example 2: Use a dictionary and function to calculate average tenure by gender
def average_tenure_by_gender(records):
    groups = {}
    for r in records:
        g = r['gender']
        if g not in groups:
            groups[g] = []
        groups[g].append(r['tenure'])
    avgs = {k: np.mean(v) for k, v in groups.items()}
    return avgs

result = average_tenure_by_gender(customers)
print(result)
{'Female': np.float64(32.24455275229358), 'Male': np.float64(32.49535864978903)}
# Intermediate Example 3: Creating a list of dictionaries for high-value customers
high_value = [r for r in customers if r['MonthlyCharges'] > 80]
print(f'Found {len(high_value)} high-value customers:')
print(high_value[:2])
Found 2666 high-value customers:
[{'gender': 'Female', 'SeniorCitizen': 0, 'Partner': 'No', 'Dependents': 'No', 'tenure': 8, 'PhoneService': 'Yes', 'MultipleLines': 'Yes', 'InternetService': 'Fiber optic', 'OnlineSecurity': 'No', 'OnlineBackup': 'No', 'DeviceProtection': 'Yes', 'TechSupport': 'No', 'StreamingTV': 'Yes', 'StreamingMovies': 'Yes', 'Contract': 'Month-to-month', 'PaperlessBilling': 'Yes', 'PaymentMethod': 'Electronic check', 'MonthlyCharges': 99.65, 'TotalCharges': '820.5', 'Churn': 'Yes'}, {'gender': 'Male', 'SeniorCitizen': 0, 'Partner': 'No', 'Dependents': 'Yes', 'tenure': 22, 'PhoneService': 'Yes', 'MultipleLines': 'Yes', 'InternetService': 'Fiber optic', 'OnlineSecurity': 'No', 'OnlineBackup': 'Yes', 'DeviceProtection': 'No', 'TechSupport': 'No', 'StreamingTV': 'Yes', 'StreamingMovies': 'No', 'Contract': 'Month-to-month', 'PaperlessBilling': 'Yes', 'PaymentMethod': 'Credit card (automatic)', 'MonthlyCharges': 89.1, 'TotalCharges': '1949.4', 'Churn': 'No'}]
# Intermediate Example 4: Define a reusable function to summarize categorical distributions
def summarize_category(records, key):
    counts = {}
    for r in records:
        v = r[key]
        if v not in counts:
            counts[v] = 1
        else:
            counts[v] += 1
    total = sum(counts.values())
    summary = {k: f'{v} ({v/total*100:.1f}%)' for k, v in counts.items()}
    return summary

contract_summary = summarize_category(customers, 'Contract')
print(contract_summary)
{'Month-to-month': '3875 (55.0%)', 'One year': '1473 (20.9%)', 'Two year': '1695 (24.1%)'}
# Intermediate Example 5: Map 'Yes'/'No' churn to 1/0 using dictionary comprehension
churn_map = {'Yes': 1, 'No': 0}
churned_numeric = [churn_map.get(r['Churn'], np.nan) for r in customers]
print(churned_numeric[:10])
[0, 0, 1, 0, 1, 1, 0, 0, 1, 0]
# Advanced Example 1: Cross-tabulate gender and churn using a nested dictionary
cross_tab = {}
for r in customers:
    g = r['gender']
    c = r['Churn']
    if g not in cross_tab:
        cross_tab[g] = {}
    if c not in cross_tab[g]:
        cross_tab[g][c] = 0
    cross_tab[g][c] += 1
print(cross_tab)
{'Female': {'No': 2549, 'Yes': 939}, 'Male': {'No': 2625, 'Yes': 930}}
# Advanced Example 2: Build a function that groups customers by a field and calculates a custom metric
def aggregate_metric(records, group_field, metric_func):
    groups = {}
    for r in records:
        key = r[group_field]
        if key not in groups:
            groups[key] = []
        groups[key].append(r)
    summary = {k: metric_func(v) for k, v in groups.items()}
    return summary

def avg_monthly_charge(records):
    return np.mean([r['MonthlyCharges'] for r in records])

charge_summary = aggregate_metric(customers, 'Contract', avg_monthly_charge)
print(charge_summary)
{'Month-to-month': np.float64(66.39849032258066), 'One year': np.float64(65.04860828241684), 'Two year': np.float64(60.77041297935103)}
# Advanced Example 3: Multi-step pipeline using lists, dictionaries, and functions to profile churn risk
def high_risk_profile(records, charge_thresh=70, max_tenure=12):
    return [r for r in records if r['MonthlyCharges'] > charge_thresh and r['tenure'] < max_tenure and r['Churn']=='Yes']

risk_customers = high_risk_profile(customers)
risk_segments = summarize_category(risk_customers, 'Contract')
print('High-churn, high-value segments:', risk_segments)
High-churn, high-value segments: {'Month-to-month': '566 (99.8%)', 'One year': '1 (0.2%)'}
# Advanced Example 4: Write and read a summary report file
with open('customer_churn_summary.txt', 'w') as f:
    for k,v in charge_summary.items():
        f.write(f'{k}: Average Monthly Charge ${v:.2f}\n')
    for contract, v in risk_segments.items():
        f.write(f'High Risk - {contract}: {v}\n')
# Error Handling Example 1: Handle missing survey responses ('TotalCharges' as blank)
num_blank = sum(1 for r in customers if r.get('TotalCharges', '') == '')
print(f'Total customers missing TotalCharges: {num_blank}')
Total customers missing TotalCharges: 0
# Error Handling Example 2: Prevent crash on incorrect group field by using dict.get
k = 'InvalidField'
try:
    value = customers[0][k]
except KeyError:
    value = None
print('Sample value:', value)
Sample value: None
# Error Handling Example 3: Detect and warn about misinterpreted NPS or Likert scales
def check_likert(nps_list):
    unexpected = [v for v in nps_list if v < 0 or v > 10]
    if unexpected:
        print(f'Warning: {len(unexpected)} NPS values are outside of 0-10 scale')
    else:
        print('All NPS scores within valid range')

np.random.seed(42)
nps_scores = list(np.random.randint(-2,13,25))
check_likert(nps_scores)
Warning: 2 NPS values are outside of 0-10 scale

Best Practices in Market Research with Lists, Dictionaries, and Functions#

  • Use clear segmentation functions to compare trends across groups.
  • For cross-tabulations, nest dictionaries for each factor (e.g., contract type by churn).
  • When making a trend index, combine counts and percents to explain both magnitude and risk.
  • For scoring, convert text responses or Yes/No into numerical indicators early.
  • Trend analysis often needs results stored in a list (for ordering) and summary dictionaries (for business reporting).
# End-to-end Market Research Problem: Predicting churn risk for monthly contract customers
monthly_customers = [r for r in customers if r['Contract'] == 'Month-to-month']
risk_metric = lambda r: int(r['MonthlyCharges'] > 75 and r['tenure'] < 6)
risks = [risk_metric(r) for r in monthly_customers]
risk_rate = sum(risks) / len(monthly_customers) * 100 if monthly_customers else 0
print(f'Monthly contract, high-risk churn rate: {risk_rate:.1f}%')
Monthly contract, high-risk churn rate: 9.4%
 

Found this useful?

All lessons, notebooks and datasets here are free. If they helped you, a coffee keeps new lessons coming.