A Python toolbox for gaining geometric insights into high-dimensional data
-
Updated
Sep 12, 2026 - Python
A Python toolbox for gaining geometric insights into high-dimensional data
📖 Use Bi-normal Separation to find document vectors which is used to compute similarity for shorter sentences.
TF-IDF sublinear vectorizer with smooth inverse document frequency and pairwise cosine similarity.
TF-IDF sublinear vectorizer with smooth inverse document frequency and pairwise cosine similarity.
Given a document, identifying the closest documents within the list of documents using tf-idf matrix and cosine similarity
Comment Sentiment Analysis using Deep Learning
🧠 Machine Learning & Natural Language Processing: Predict the author of literary text snippets. Built with TensorFlow and Keras, this project trains an LSTM model on classic literature to identify writing style and authorship.
Machine learning-based sentiment analysis system that classifies movie reviews as positive or negative using NLP techniques and text classification models.
A simple Python script for transforming a corpus of documents into text vectors suitable for visualization
AI-powered Book Recommendation System using Django, TF-IDF, and semantic search (RAG-based).
Resume Matcher: A Streamlit app that compares resumes with job descriptions, highlights matched and missing skills, and calculates text similarity and skill coverage.
Homeworks and final project for Infosearch course
This program is a project carried out in the Natural Language Processing course, which is a Taylor Swift song recommender. It utilizes topics such as sentiment analysis in texts, text vectorization, and the removal of stopwords.
Naive Bayes classifier with text parser and vectorization libs
Machine Learning course of Piero Savastano 3: CountVectorizer
AI Resume & Job Matching Tool is a Streamlit-based web app that compares a candidate’s resume with a job description using NLP (TF-IDF & cosine similarity). It generates a match score and highlights missing skills, helping job seekers optimize their resumes for specific roles.
The project finds the emotion in the given text. It takes input in a text/sentence format through the User or through Twitter API.
An end-to-end NLP pipeline for detecting hate speech in tweets using TF-IDF and Logistic Regression, with hyperparameter tuning via Grid Search. Features a Streamlit web UI with lexical explainability and an agentic LLM layer (LangChain + LLaMA 3) for nuanced policy violation analysis of flagged content.
An end-to-end Transformer-based sentiment classification system built with TensorFlow/Keras, demonstrating text preprocessing, tokenization, self-attention, sequence encoding, and binary text classification using an Encoder-only Transformer architecture.
To associate your repository with the text-vectorization topic, visit your repo's landing page and select "manage topics."