BANK MARKETING STUDY · 2026

Term Deposit
Subscription
Prediction

Comparing Logistic Regression and Random Forest to identify customers most likely to respond to a direct-marketing campaign.

R 4.5.2caretrandomForestpROCPCAVIF

RESEARCH OVERVIEW

Who is likely to subscribe?

The study predicts whether a customer will subscribe to a bank term deposit using demographic, financial, contact, economic, and previous-campaign information. It contrasts an interpretable parametric model with a flexible ensemble learner.

RQ

Research questionWhich model best identifies likely subscribers while producing useful evidence for campaign decisions?

01 · DATASET

Prepared for realistic pre-call prediction.

Observations41,188Portuguese bank campaigns
Original variables2120 predictors after leakage removal
Subscribers11.27%Strong minority class
Missing values0Unknown retained as a category

Target distribution

Most customers did not subscribe.

Accuracy alone is misleading when 88.73% belong to the “no” class, so recall, precision, F1, and ROC-AUC were essential.
88.73% No subscriptionMajority
11.27% SubscribedTarget

02 · EXPLORATORY ANALYSIS

Campaign response is shaped by several signals.

01

Duration removed

Leakage

Call duration is only known after contact and would create unrealistic performance.

02

Repeated contact skews

Right

Most customers were contacted only a few times, suggesting diminishing returns.

03

Economic overlap

High VIF

Euribor, employment variation, and employee count showed multicollinearity.

04

Compact structure

70%

The first three principal components explained about 70% of numeric variance.

Five predictor families

DemographicAge · job · education
FinancialDefault · housing · loan
CampaignContact · month · attempts

Important campaign signals

Previous outcomepoutcome
Economic climateemp_var_rate
Timing and channelmonth · contact

03 · METHODOLOGY

A leakage-safe modelling pipeline.

01CleanStandardise names; retain unknown levels
02TransformFactors and Z-score numeric scaling
03Split70% train · 30% test · stratified
04ModelTwo logistic and two forest variants
05EvaluateAccuracy, precision, recall, F1, AUC
INTERPRETABILITY

Logistic Regression

Coefficients and odds ratios translate directly into campaign guidance, with VIF checks supporting responsible interpretation.

NON-LINEARITY

Random Forest

Tree ensembles model interactions and rank variable importance without relying on normal distributions.

CLASS IMBALANCE

Recall matters

Finding genuine subscribers has greater campaign value than simply predicting the majority class correctly.

04 · MODEL PERFORMANCE

Interpretability versus minority-class detection.

ModelAccuracyPrecisionRecallF1AUCTreesBest use
Logistic Regression 1 BEST AUC.9004.6695.2284.3407.7888Interpretation
Logistic Regression 2.8996.6624.2227.3333.7884Simplified
Random Forest 1.8971.5872.2902.3885.7783300Recall
Random Forest 2 BEST F1.8979.5948.2931.3927.7799500Targeting

05 · INTERPRETATION

Previous outcomes and economic context matter.

Previous outcome100
Employment rate86
Consumer prices78
Contact month72
Contact channel64
Euribor 3m59
Campaign contacts48
BUSINESS FINDING

Model choice depends on campaign goals.

Logistic Regression offered the strongest AUC and clearer managerial explanation. The tuned Random Forest found more actual subscribers and achieved the best F1 score, making it more useful where positive-case identification is the priority.

High accuracy does not guarantee an effective marketing model when the valuable response class is rare.

06 · CONCLUSION

Use interpretable probability and class-aware targeting together.

Logistic Regression achieved the best overall discrimination at AUC 0.7888, while the 500-tree Random Forest improved recall to 0.2931 and F1 to 0.3927.

The analysis demonstrates why leakage prevention, stratified evaluation, and business-relevant metrics are essential when building real-world campaign models.

Back to data science projects →