23/07/2026
Excited to share our latest project at Codify Labs — a Smart Inbox Spam Classifier built using Machine Learning. This project was assigned to all interns (individually) as part of their academic assignment, and they have successfully designed, trained, and deployed it end-to-end.
The team trained and compared six different models before finalizing the best one:
1) Multinomial Naive Bayes
2) Gaussian Naive Bayes
3) Bernoulli Naive Bayes
4) Logistic Regression
5) SVM (Linear)
6) Random Forest
Performance comparison (top models):
Multinomial Naive Bayes — Accuracy: 98.45%, Precision: 98.60%, Recall: 98.40%, F1-Score: 98.49%
Logistic Regression — Accuracy: 97.60%, Precision: 97.70%, Recall: 97.40%, F1-Score: 97.54%
SVM (Linear) — Accuracy: 96.20%, Precision: 96.30%, Recall: 96.10%, F1-Score: 96.19%
After evaluating all six models, Multinomial Naive Bayes was selected as the final model for its highest accuracy and the most balanced performance across all metrics.
Technical Details:
Dataset: Public SMS Spam Collection Dataset (5,574 messages)
Preprocessing: Lowercasing, punctuation and special character removal, stop-word removal, tokenization
Feature Extraction: TF-IDF Vectorization (max_features=5000)
Tools and Libraries: Python, Scikit-learn, Pandas, NumPy, NLTK, Streamlit, Matplotlib, Seaborn
Results:
98.45% accuracy
Tested on 1,250+ real messages
820 correctly flagged as Spam, 430 as Ham (Safe)
Key Findings:
Spam messages were detected with high precision (98.60%) and recall (98.40%)
TF-IDF features played a crucial role in capturing important patterns in text data
The model generalizes well on unseen data and is lightweight for real-time prediction
A responsive web app was built using Streamlit for real-time user predictions and analytics
This project reflects our interns' strong grasp of Machine Learning fundamentals, from model selection to deployment.
Building smart solutions, one model at a time.