Project Business Statistics: E-news Express¶
Define Problem Statement and Objectives¶
Import all the necessary libraries¶
# Installing the libraries with the specified version.
!pip install numpy==1.25.2 pandas==1.5.3 matplotlib==3.7.1 seaborn==0.13.1 scipy==1.11.4 -q --user
import numpy as np
import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt
%matplotlib inline
import scipy.stats as stats
from statsmodels.stats.proportion import proportions_ztest
import warnings
warnings.simplefilter('ignore')
pd.set_option('display.float_format', lambda x: '%.2f' % x)
sns.set_palette("tab10")
Note: After running the above cell, kindly restart the notebook kernel and run all cells sequentially from the start again.
Business Context¶
The advent of e-news, or electronic news, portals has offered us a great opportunity to quickly get updates on the day-to-day events occurring globally. The information on these portals is retrieved electronically from online databases, processed using a variety of software, and then transmitted to the users. There are multiple advantages of transmitting new electronically, like faster access to the content and the ability to utilize different technologies such as audio, graphics, video, and other interactive elements that are either not being used or aren’t common yet in traditional newspapers.
E-news Express, an online news portal, aims to expand its business by acquiring new subscribers. With every visitor to the website taking certain actions based on their interest, the company plans to analyze these actions to understand user interests and determine how to drive better engagement. The executives at E-news Express are of the opinion that there has been a decline in new monthly subscribers compared to the past year because the current webpage is not designed well enough in terms of the outline & recommended content to keep customers engaged long enough to make a decision to subscribe.
Objective:¶
The design team of the company has researched and created a new landing page that has a new outline & more relevant content shown compared to the old page. In order to test the effectiveness of the new landing page in gathering new subscribers, the Data Science team conducted an experiment by randomly selecting 100 users and dividing them equally into two groups. The existing landing page was served to the first group (control group) and the new landing page to the second group (treatment group). Data regarding the interaction of users in both groups with the two versions of the landing page was collected. Being a data scientist in E-news Express, you have been asked to explore the data and perform a statistical analysis (at a significance level of 5%) to determine the effectiveness of the new landing page in gathering new subscribers for the news portal by answering the following questions:
- Do the users spend more time on the new landing page than on the existing landing page?
- Is the conversion rate (the proportion of users who visit the landing page and get converted) for the new page greater than the conversion rate for the old page?
- Does the converted status depend on the preferred language?
- Is the time spent on the new page the same for the different language users?
Data Dictionary¶
The data contains information regarding the interaction of users in both groups with the two versions of the landing page.
- user_id - Unique user ID of the person visiting the website
- group - Whether the user belongs to the first group (control) or the second group (treatment)
- landing_page - Whether the landing page is new or old
- time_spent_on_the_page - Time (in minutes) spent by the user on the landing page
- converted - Whether the user gets converted to a subscriber of the news portal or not
- language_preferred - Language chosen by the user to view the landing page
Reading the Data into a DataFrame¶
df = pd.read_csv('abtest.csv')
Exploratory Data Analysis¶
Dataset Information and Structure¶
df.shape
(100, 6)
df.info()
<class 'pandas.core.frame.DataFrame'> RangeIndex: 100 entries, 0 to 99 Data columns (total 6 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 user_id 100 non-null int64 1 group 100 non-null object 2 landing_page 100 non-null object 3 time_spent_on_the_page 100 non-null float64 4 converted 100 non-null object 5 language_preferred 100 non-null object dtypes: float64(1), int64(1), object(4) memory usage: 4.8+ KB
Random sample of data¶
df.sample(5)
| user_id | group | landing_page | time_spent_on_the_page | converted | language_preferred | |
|---|---|---|---|---|---|---|
| 94 | 546550 | control | old | 3.05 | no | English |
| 49 | 546473 | treatment | new | 10.50 | yes | English |
| 15 | 546466 | treatment | new | 6.27 | yes | Spanish |
| 35 | 546552 | control | old | 8.50 | yes | English |
| 67 | 546582 | control | old | 4.75 | yes | Spanish |
Summary Statistics of the Dataset¶
df.describe()
| user_id | time_spent_on_the_page | |
|---|---|---|
| count | 100.00 | 100.00 |
| mean | 546517.00 | 5.38 |
| std | 52.30 | 2.38 |
| min | 546443.00 | 0.19 |
| 25% | 546467.75 | 3.88 |
| 50% | 546492.50 | 5.42 |
| 75% | 546567.25 | 7.02 |
| max | 546592.00 | 10.71 |
Check for missing data¶
df.isna().sum()
user_id 0 group 0 landing_page 0 time_spent_on_the_page 0 converted 0 language_preferred 0 dtype: int64
Crosstab for Language Preferred and Converted Status¶
crosstab_language_converted = pd.crosstab(df['language_preferred'], df['converted'], margins=True, normalize='index')
crosstab_language_converted
| converted | no | yes |
|---|---|---|
| language_preferred | ||
| English | 0.34 | 0.66 |
| French | 0.56 | 0.44 |
| Spanish | 0.47 | 0.53 |
| All | 0.46 | 0.54 |
Crosstab for Landing Page and Converted¶
crosstab_group_converted = pd.crosstab(df['landing_page'], df['converted'], margins=True, normalize='index')
crosstab_group_converted
| converted | no | yes |
|---|---|---|
| landing_page | ||
| new | 0.34 | 0.66 |
| old | 0.58 | 0.42 |
| All | 0.46 | 0.54 |
Percentage increase calculations¶
conv_rate_old = df[df['landing_page'] == 'old']['converted'].value_counts()['yes']
conv_rate_new = df[df['landing_page'] == 'new']['converted'].value_counts()['yes']
total_conversion_percentage_increase = ((conv_rate_new - conv_rate_old) / conv_rate_old) * 100
time_old = df[df['landing_page'] == 'old']['time_spent_on_the_page'].mean()
time_new = df[df['landing_page'] == 'new']['time_spent_on_the_page'].mean()
total_time_percentage_increase = ((time_new - time_old) / time_old) * 100
languages = df['language_preferred'].unique()
conversion_percentage_increase_language = {}
time_spent_percentage_increase_language = {}
for language in languages:
conversion_rate_old_lang = df[(df['landing_page'] == 'old') & (df['language_preferred'] == language)]['converted'].value_counts().get('yes', 0)
conversion_rate_new_lang = df[(df['landing_page'] == 'new') & (df['language_preferred'] == language)]['converted'].value_counts().get('yes', 0)
conversion_percentage_increase_language[language] = ((conversion_rate_new_lang - conversion_rate_old_lang) / conversion_rate_old_lang) * 100
time_spent_old_lang = df[(df['landing_page'] == 'old') & (df['language_preferred'] == language)]['time_spent_on_the_page'].mean()
time_spent_new_lang = df[(df['landing_page'] == 'new') & (df['language_preferred'] == language)]['time_spent_on_the_page'].mean()
time_spent_percentage_increase_language[language] = ((time_spent_new_lang - time_spent_old_lang) / time_spent_old_lang) * 100
print('Total percentage increase for new landing page:', total_conversion_percentage_increase)
print('Total time spent percentage increase for new landing page:', total_time_percentage_increase)
print()
print('Conversion percentage increase for Spanish:', conversion_percentage_increase_language['Spanish'])
print('Conversion percentage increase for English:', conversion_percentage_increase_language['English'])
print('Conversion percentage increase for French:', conversion_percentage_increase_language['French'])
print()
print('Time spent percentage increase for Spanish:', time_spent_percentage_increase_language['Spanish'])
print('Time spent percentage increase for English:', time_spent_percentage_increase_language['English'])
print('Time spent percentage increase for French:', time_spent_percentage_increase_language['French'])
Total percentage increase for new landing page: 57.14285714285714 Total time spent percentage increase for new landing page: 37.30473921101401 Conversion percentage increase for Spanish: 57.14285714285714 Conversion percentage increase for English: -9.090909090909092 Conversion percentage increase for French: 300.0 Time spent percentage increase for Spanish: 20.85769980506822 Time spent percentage increase for English: 49.600112249193224 Time spent percentage increase for French: 43.76961921659619
Perform Visual Analysis¶
plt.figure(figsize=(18, 12))
# Box plots
plt.subplot(2, 3, 1)
sns.boxplot(y='time_spent_on_the_page', data=df)
plt.title('Time Spent on Page (New and Old)')
plt.subplot(2, 3, 2)
sns.boxplot(x='landing_page', y='time_spent_on_the_page', data=df)
plt.title('Time Spent by Landing Page')
plt.subplot(2, 3, 3)
sns.boxplot(x='converted', y='time_spent_on_the_page', data=df)
plt.title('Time Spent by Conversion Status')
# Histograms
ax1 = plt.subplot(2, 3, 4)
sns.histplot(df['time_spent_on_the_page'], kde=True, ax=ax1)
plt.title('Time Spent on Page (New and Old)')
ax2 = plt.subplot(2, 3, 5)
sns.histplot(data=df, x='time_spent_on_the_page', hue='landing_page', element='step', kde=True, ax=ax2)
plt.title('Time Spent by Landing Page')
ax3 = plt.subplot(2, 3, 6)
sns.histplot(data=df, x='time_spent_on_the_page', hue='converted', element='step', kde=True, ax=ax3)
plt.title('Time Spent by Conversion Status')
# Get the maximum y-value from all three plots
max_y = max(ax1.get_ylim()[1], ax2.get_ylim()[1], ax3.get_ylim()[1])
# Set the same y-axis limit for all subplots
ax1.set_ylim(0, max_y)
ax2.set_ylim(0, max_y)
ax3.set_ylim(0, max_y)
plt.tight_layout()
plt.show()
Observations:
- There appears to be an increase in time spent and conversion status on the new landing page.
# Box plots for time spent on the page by language, divided by landing page (new and old)
plt.figure(figsize=(10, 6))
sns.boxplot(x='language_preferred', y='time_spent_on_the_page', hue='landing_page', data=df)
plt.title('Time Spent on the Page by Language and Landing Page')
plt.show()
Observations:
- These box plots show an increase in time spent on the new page in total and across all languages.
# Set the desired order for the hue categories
hue_order = ['no', 'yes']
plt.figure(figsize=(16, 5))
# Bar plots
sns.countplot(x='landing_page', hue='converted', data=df, hue_order=hue_order)
plt.title("Conversion by Landing Page and Status", fontsize=16)
plt.xlabel("Landing Page", fontsize=12)
plt.ylabel("Count", fontsize=12)
plt.legend(title='Converted', loc='upper right')
plt.show()
# Facet grid of bar plots by language with conversion status
g = sns.FacetGrid(df, col="language_preferred", height=5, aspect=1, margin_titles=True)
g.map_dataframe(sns.countplot, x="landing_page", hue="converted", hue_order=hue_order, palette='tab10')
# Add a legend and title
g.set_axis_labels("Landing Page", "Count")
g.set_titles(col_template="{col_name}", size=16, weight='bold')
# Add borders to each subplot
for ax in g.axes.flatten():
ax.axhline(y=ax.get_ylim()[1], color='black', linestyle='-')
ax.axhline(y=ax.get_ylim()[0], color='black', linestyle='-')
ax.axvline(x=ax.get_xlim()[0], color='black', linestyle='-')
ax.axvline(x=ax.get_xlim()[1], color='black', linestyle='-')
# Adjust legend position within each row's dimensions
for ax in g.axes.flatten():
legend = ax.legend(loc='upper right')
legend.get_frame().set_edgecolor('black')
plt.suptitle("Conversion by Landing Page, Status, and Language", y=1.05, fontsize=16)
plt.show()
Observations:
- These bar charts show an increase in the conversion rates with the new page in total and across all languages.
contingency_table = pd.crosstab(df['landing_page'], df['converted'])
plt.figure(figsize=(8, 6))
sns.heatmap(contingency_table, annot=True, fmt='d', cmap='Blues')
plt.title('Landing Page vs. Converted')
plt.xlabel('Converted')
plt.ylabel('Landing Page')
plt.show()
Observations:
- The table shows a strong relationship between the new page and higher conversion rate.
Perform Hypothesis Testing¶
1. Do the users spend more time on the new landing page than the existing landing page?¶
plt.figure(figsize=(10,1.5))
meanprops = {'marker': '.','markerfacecolor': 'white','markeredgecolor': 'white','markersize': 5}
sns.boxplot(y = 'landing_page', x = 'time_spent_on_the_page', showmeans=True, meanprops=meanprops, data = df)
plt.show()
The null and alternative hypothesis:
$H_0:$ $μ_{\text{new}}$ ≤ $μ_{\text{old}}$ - The mean time spent on the new landing page is less than or equal to the mean time spent on the existing landing page.
$H_a:$ $μ_{\text{new}}$ > $μ_{\text{old}}$ - The mean time spent on the new landing page is greater than the mean time spent on the existing landing page.
To compare the sample means from two independent populations where the standard deviation is unknown the 2-sample independent t-test is the best choice.
The significance level is 0.05%.
#Collect and prepare data
new_page_time = df[df['landing_page'] == 'new']['time_spent_on_the_page']
old_page_time = df[df['landing_page'] == 'old']['time_spent_on_the_page']
#Calculate the p-value
t_stat, p_value = stats.ttest_ind(new_page_time, old_page_time, alternative='greater')
print(f"p-value: {p_value}")
p-value: 0.00013161235280950055
The p-value of 0.00013 < 0.05% significance level
Since the p-value of 0.00013 is lower than the significanse level of 0.05% we reject the null hypothesis.
In conclusion: We reject the null hypothesis and conclude that users spend significantly more time on the new landing page compared to the existing landing page.
2. Is the conversion rate (the proportion of users who visit the landing page and get converted) for the new page greater than the conversion rate for the old page?¶
pd.crosstab(df['landing_page'],df['converted'],normalize='index').plot(kind="barh", figsize=(6,2), stacked=False)
plt.legend(loc='upper center', bbox_to_anchor=(0.17, 1.3), ncol=2)
plt.show()
The null and alternative hypothesis:
$H_0:$ $p_{\text{new}}$ ≤ $p_{\text{old}}$ - The conversion rate for the new landing page is less than or equal to the conversion rate for the existing landing page.
$H_a:$ $p_{\text{new}}$ > $p_{\text{old}}$ - The conversion rate for the new landing page is greater than the conversion rate for the existing landing page.
To compare the sample proportions from two populations the 2 sample z-test is the best choice.
The significance level is 0.05%.
#Collect and prepare data
conversions_new = df[(df['landing_page'] == 'new') & (df['converted'] == 'yes')].shape[0]
conversions_old = df[(df['landing_page'] == 'old') & (df['converted'] == 'yes')].shape[0]
total_new = df[df['landing_page'] == 'new'].shape[0]
total_old = df[df['landing_page'] == 'old'].shape[0]
count = np.array([conversions_new, conversions_old])
nobs = np.array([total_new, total_old])
print('Conversion counts new, old:', count)
print('Number of observations new, old: ', nobs)
Conversion counts new, old: [33 21] Number of observations new, old: [50 50]
#Calculate the p-value
z_stat, p_value = proportions_ztest(count, nobs, alternative='larger')
print(f"p-value: {p_value}")
p-value: 0.008026308204056278
The p-value of 0.008 < 0.05% significance level
Since the p-value of 0.008 is lower than the significanse level of 0.05% we reject the null hypothesis.
Based on this test, we conclude that the new landing page does have a larger conversion rate than the old landing page.
3. Is the conversion and preferred language are independent or related?¶
pd.crosstab(df['converted'],df['language_preferred'],normalize='index').plot(kind="barh", figsize=(6,3), stacked=False)
plt.legend(loc='upper center', bbox_to_anchor=(0.36, 1.2), ncol=3)
plt.show()
The null and alternative hypothesis:
$H_0:$ The conversion rate and language are independent
$H_a:$ The conversion rate depends on language
To determine if the conversion rate is independent from language the Chi-square test ofor independence is the best choice.
The significance level is 0.05%.
#Collect and prepare data
contingency_table = pd.crosstab(df['converted'], df['language_preferred'])
contingency_table
| language_preferred | English | French | Spanish |
|---|---|---|---|
| converted | |||
| no | 11 | 19 | 16 |
| yes | 21 | 15 | 18 |
#Calculate the p-value
chi2_stat, p_value, dof, expected = stats.chi2_contingency(contingency_table)
print(f"p-value: {p_value}")
p-value: 0.21298887487543447
The p-value of 0.213 > 0.05% significance level
Since the p-value of 0.213 is higher than the significanse level of 0.05% we fail to reject the null hypothesis.
Based on this test, there is not enough evidence to suggest that the conversion rate is dependent on the language.
4. Is the time spent on the new page same for the different language users?¶
df_new = df[df['landing_page'] == 'new']
df_mean_by_language_preferred = df_new.groupby(['language_preferred'])['time_spent_on_the_page'].mean()
df_mean_by_language_preferred
language_preferred English 6.66 French 6.20 Spanish 5.84 Name: time_spent_on_the_page, dtype: float64
plt.figure(figsize=(10,2))
meanprops = {'marker': '.','markerfacecolor': 'white','markeredgecolor': 'white','markersize': 5}
ax = sns.boxplot(y = 'language_preferred', x = 'time_spent_on_the_page', showmeans = True,meanprops=meanprops, data = df_new)
plt.show()
The null and alternative hypothesis:
$H_0:$ $μ_{\text{English}}$ = $μ_{\text{French}}$ = $μ_{\text{Spanish}}$
$H_a:$ The μ time spent on the new landing page is different for at least one language.
To compare the means of time spent on the new landing page of 3 independent language populations the ANOVA test is the best choice. __
__The significance level is 0.05%.
#Collect and prepare data
new_page_data = df[df['landing_page'] == 'new']
english_time = new_page_data[new_page_data['language_preferred'] == 'English']['time_spent_on_the_page']
french_time = new_page_data[new_page_data['language_preferred'] == 'French']['time_spent_on_the_page']
spanish_time = new_page_data[new_page_data['language_preferred'] == 'Spanish']['time_spent_on_the_page']
#Calculate the p-value
f_stat, p_value = stats.f_oneway(english_time, french_time, spanish_time)
print(f"p-value: {p_value}")
p-value: 0.43204138694325955
The p-value of 0.432 > 0.05% significance level
Since the p-value of 0.432 is higher than the significanse level of 0.05% we fail to reject the null hypothesis.
Based on this test, there is not enough evidence to suggest that the average time spent on the new landing page is different for any of the languages.
Conclusion and Business Recommendations¶
- Users spend more time on the new landing page than the old landing page.
- The conversion rate for the new landing page is higher than the old landing page.
- The higher average time spent and conversion rates for the new page are independent of the language.
- It is recommended that all new users get directed to the new landing page.
In general there is an increase in conversion rate and time spent on the new landing page but there is one exception with the new page in English there is a 9% decrease in conversion even though there was still an increase in time spent. Additionally there is a significantly higher conversion rate for French.
- Total percentage increase for new landing page: 57%
- Total time spent percentage increase for new landing page: 37%
- Conversion percentage increase for Spanish: 57%
- Conversion percentage increase for English: -9%
- Conversion percentage increase for French: 300%
- Time spent percentage increase for Spanish: 20%
- Time spent percentage increase for English: 49%
- Time spent percentage increase for French: 43%
It is recommended to:
- Take a closer look at the content on the English page to determine what could be causing the decrease in conversion.
- Take a closer look at the content on the French page to determine what could be causeing such a high conversion rate.
- Determine new changes to test based on analysis from the new French and English pages.
- Run another AB test with the proposed change and reanalyze the data the next iteration of changes.