Evalica, your favourite evaluation toolkit
-
Updated
Sep 29, 2026 - Python
Evalica, your favourite evaluation toolkit
GEditBench v2: A Human-Aligned Benchmark for General Image Editing
A visual complexity dataset across seven different categories, including Scenes, Advertisements, Visualization and infographics, Objects, Interior design, Art, and Suprematism for computer vision application.
Compare Elo, Glicko, TrueSkill, Bradley-Terry, and other rating algorithms behind a largely uniform Python interface.
An interactive web application for comparing entities using the pairwise comparison method
sort by meaning: order lines along a plain-English dimension, from pairwise comparisons judged by TypeSafe's Jev model
An interactive Shiny App for UpSet plots, Venn diagrams and Pairwise heatmaps
🏫 🇧🇩 Fuzzy-AHP-based recommendation system for secondary schools in Bangladesh
Concept-Guided Chain-of-Thought (CGCoT) pairwise annotation tool for systematic text evaluation using LLMs. Generate breakdowns, compare items, compute scores, and validate against human judgments. Supports Ollama, Hugging Face, Google Gemini, OpenAI, and Anthropic models.
Fast, large scale library for computing rankings and features based on various pairwise and graph algorithms
Predicting missing pairwise preferences from similarity features in group decision making and group recommendation system
A Jupyter notebook for a project centered around 'Group Recommendation Systems (GRS)' utilizing the 'GcPp' clustering approach.
AI-powered personal recommendation engine that learns preferences through pairwise comparison
Framework for using LLMs to grade texts by using pairwise comparisons.
Adversarial Preference Learning with Pairwise Comparisons for Group recommendation System
Web-based decision tool implementing the Analytic Hierarchy Process (AHP): break complex choices into pairwise comparisons, check your consistency, and get a transparent, repeatable ranking — no spreadsheets required.
A personality-aware group recommendation system based on pairwise preferences
Sampling algorithm for best-worst scaling sets.
Learning to rank (LTR) with noisy labels and simulation of how noise affects quality of ranking
To associate your repository with the pairwise-comparison topic, visit your repo's landing page and select "manage topics."