Six-model comparison
Logistic Regression, Linear SVM, XGBoost, LSTM, 1D CNN and DistilBERT are compared in the study.
Project 04 / 12 · NLP / Machine Learning
A research-paper classification study and interactive application comparing six models across six academic categories.

ResearchScope AI was developed by Group 20 — ROT NLP Solutions. The study compares six machine-learning and deep-learning models using a balanced dataset of 15,000 arXiv papers across six academic categories.
As Team Lead, my work covered preprocessing, TF-IDF, Logistic Regression and LSTM. The best-performing model, Logistic Regression, achieved 89.33% test accuracy in the reported evaluation.
The Streamlit application makes the research interactive: users can explore predictions, confidence scores, the top three categories, keyword insights and model comparisons. The work ran from June to August 2026.
Logistic Regression, Linear SVM, XGBoost, LSTM, 1D CNN and DistilBERT are compared in the study.
15,000 arXiv papers provide a balanced basis for evaluation across six categories.
The application exposes confidence scores, top-three predictions and keyword insights.
Logistic Regression achieved 89.33% test accuracy within the project’s evaluation setup.
Prepare → Clean and preprocess the research-paper text for model input.
Represent & train → Build TF-IDF representations and train classical and neural approaches.
Evaluate & demonstrate → Compare the models and expose predictions through Streamlit.
Open an image to see the full detail. Use the arrow keys to move through the gallery.
Official references for the tools used in this project.