--- title: BenchHub Leaderboards emoji: 🏆 colorFrom: purple colorTo: indigo sdk: static app_file: index.html pinned: true hf_oauth: false short_description: Reproducible LLM/vision/audio leaderboards. Submit free tags: - leaderboard - benchmark - evaluation - model-evaluation - reproducibility - llm - text-generation - reasoning - mmlu - computer-vision - image-classification - object-detection - image-segmentation - depth-estimation - optical-flow - stereo - audio-classification - automatic-speech-recognition - speech-recognition - question-answering - named-entity-recognition - token-classification datasets: - AI-Lab-Makerere/beans - Aaryan333/fer2013_train_publicTest_privateTest - Bingsu/Cat_and_Dog - Bingsu/Human_Action_Recognition - Falah/Alzheimer_MRI - Hemg/Brain-Tumor-MRI-Dataset - Joker5800/Tobacco_Leaf_Diseases - Marxulia/asl_sign_languages_alphabets_v03 - Multimodal-Fatima/Caltech101_not_background_test - Multimodal-Fatima/FGVC_Aircraft_test - Nattakarn/fruit-and-vegetable-image-recognition - POSE-Lab/IndustryShapes - Rishi210904/Crop_Weed_classification - Rowan/hellaswag - Spitblaze/Corn_or_Maize_Leaf_Disease_Dataset - TIGER-Lab/MMLU-Pro - Teklia/IAM-line - TheKernel01/140k-Real-and-Fake-Faces - abisee/cnn_dailymail - ahmed-ai/skin-lesions-classification-dataset - allenai/ai2_arc - allenai/openbookqa - allenai/qasc - allenai/sciq - allenai/winogrande - amaye15/stanford-dogs - anthony2261/paddy-disease-classification - aps/super_glue - ashraq/esc50 - blanchon/EuroSAT_RGB - blanchon/PatternNet - blanchon/UC_Merced - cais/mmlu - cornell-movie-review-data/rotten_tomatoes - dair-ai/emotion - deanngkl/raf-db-7emotions - deepmind/aqua_rat - detection-datasets/fashionpedia - dpdl-benchmark/oxford_flowers102 - efekankavalci/CUB_200_2011 - ehovy/race - ethz/food101 - facebook/anli - fancyzhx/ag_news - flaviagiammarino/vqa-rad - google-research-datasets/conceptual_captions - google/boolq - google/speech_commands - hails/agieval-logiqa-en - hails/agieval-lsat-lr - jagennath-hari/nyuv2 - jonathan-roberts1/GID - jonathan-roberts1/NWPU-RESISC45 - jonathan-roberts1/Optimal-31 - jonathan-roberts1/SIRI-WHU - jonathan-roberts1/Satellite-Images-of-Hurricane-Damage - jonathan-roberts1/USTC_SmokeRS - keremberke/blood-cell-object-detection - keremberke/csgo-object-detection - keremberke/excavator-detector - keremberke/forklift-object-detection - keremberke/hard-hat-detection - keremberke/license-plate-object-detection - keremberke/plane-detection - keremberke/valorant-object-detection - lighteval/piqa - lighteval/siqa - liventruth/PSAS-Acoustic-Verification - naufalso/carla_hd - nelorth/oxford-flowers - openai/gsm8k - openlifescienceai/medmcqa - openslr/librispeech_asr - prithivMLmods/Face-Age-10K - prs-eth/ZuriPano - qingyuyang/Fetal_Planes_DB - rafaelpadilla/coco2017 - rajpurkar/squad - rotemvahava/yoga-poses-107 - sartajbhuvaji/Brain-Tumor-Classification - stanfordnlp/imdb - stanfordnlp/snli - stanfordnlp/sst2 - stereo-dataset/stereo-dataset - tanganke/dtd - tanganke/eurosat - tanganke/gtsrb - tanganke/kmnist - tanganke/resisc45 - tanganke/stanford_cars - tanganke/stl10 - tau/commonsense_qa - timm/mini-imagenet - timm/oxford-iiit-pet - timm/resisc45 - trpakov/chest-xray-classification - truthfulqa/truthful_qa - uoft-cs/cifar10 - uoft-cs/cifar100 - wmt/wmt14 - xbgoose/ravdess - ylecun/mnist - zalando-datasets/fashion_mnist - zh-plus/tiny-imagenet --- # 🏆 BenchHub Leaderboards A read-only mirror of **public** leaderboard standings from **[runbenchhub.com](https://runbenchhub.com)** — an open, multi-modal model benchmarking platform. Browse the standings here; **run your own model and submit on BenchHub** (free). ## What's benchmarked Live boards across a growing set of domains, each with real models scored on the same eval set: - **LLM / Reasoning** — MMLU (14k MCQ, 57 subjects), HellaSwag, ARC-Challenge, GSM8K — one **pinned prompt**, scored by exact match, zero-shot, so every model is evaluated identically - **Vision** — Image Classification, Semantic Segmentation (mIoU), Object Detection (mAP), Optical Flow, Monocular & **Stereo** Depth Estimation, Point Tracking, Image Captioning - **Audio** — Audio Classification (ESC-50), Automatic Speech Recognition (WER) - **NLP** — Text Classification, Extractive Question Answering (SQuAD), Named-Entity Recognition (CoNLL-2003), Visual Question Answering, Translation ## Submit your model Every **"Submit"** button here deep-links to that leaderboard's submission page on BenchHub. Submitting is free — sign in with **GitHub, Google, or 🤗 Hugging Face**, run the one-line client on your predictions, and see where you rank. > This Space holds **no** ground-truth samples, predictions, or submission UI — > it reads a derived standings dataset (`HF_RESULTS_REPO`) that BenchHub > publishes. The full interactive experience (per-sample explorer, GT > visualizations, model comparison) lives on > **[runbenchhub.com](https://runbenchhub.com)**.