-
Benchmarks & evals Open arena for image-to-3D generation on Hugging Face, running since June 2024: voters compare anonymized 3D models side by side, and more than 120,000 pairwise votes feed an Elo leaderboard.
-
APIsBusiness Job-postings data — the API behind many salary and vacancy trackers.
-
Benchmarks & evalsAcademic Berkeley RDI's large-scale agent benchmark: long-horizon, economically valuable professional tasks spanning 55 sub-industries.
-
Responsible AISocial science The Responsible AI Collaborative's crowdsourced catalog of real-world AI harms: 1,600+ incidents, each linked to the news reports behind it and tagged with the CSET, GMF, and MIT taxonomies, with weekly snapshot downloads in JSON, CSV, and MongoDB formats.
-
Security & privacy Open community database of failures in general-purpose AI systems, from jailbreaks and prompt injection to vulnerabilities in AI tooling, split into evidence-backed reports and recurring vulnerability records, with a Python SDK.
-
BusinessSocial science Anthropic's periodic data releases on how its Claude models are used across the economy: privacy-preserving analysis of conversations mapped to O*NET occupational tasks and split by automation versus augmentation, country, and US state. Datasets on Hugging Face since February 2025.
-
APIs Web-scraping platform with a marketplace of ready-made extractors.
-
Benchmarks & evals The ARC-AGI abstraction-and-reasoning benchmark and prize: public training and eval sets, a leaderboard tracking progress toward general intelligence, and ARC-AGI-3's interactive game environments that test learning without weight updates.
-
GeospatialAggregated lists Global search across every ArcGIS open-data hub.
-
Benchmarks & evalsAggregated lists Independent, continuously refreshed comparisons of AI models and providers: intelligence, speed, and price on one axis system, plus blind-vote arenas for image, video, speech, and music generation (its Music Arena is separate from Carnegie Mellon's).
-
EnergyGovernment
-
Benchmarks & evals Microsoft's benchmark for text-to-audio-video generation, the models that produce sound with the picture: scores audio and visual quality, lip and sound synchronization, physical plausibility, and control over speech, music, and on-screen text. Results table on the site, dataset on Hugging Face.
-
Synthetic dataBusiness Privacy-preserving data collaboration, including synthetic dataset generation.
-
Aggregated lists
-
GovernmentCensus
-
Benchmarks & evalsAggregated lists Aggregated LLM leaderboard and model-comparison site rolling many published benchmarks into one browsable view.
-
GovernmentCensus
-
Benchmarks & evals ClickHouse's open benchmark of ~50 analytical databases on a shared web-analytics workload; methodology and raw results public.
-
Social science
-
Benchmarks & evalsAcademic Expert-validated tasks for agents that learn across sequences of task instances instead of one-off completions; open codebase and paper.
-
GovernmentSocial science
-
HealthAggregated lists A catalog of 6,000+ biological databases worldwide.
-
Benchmarks & evals Datacurve's contamination-free coding-agent benchmark: 113 original long-horizon tasks written from scratch across 91 active repositories and five languages, graded by hand-written functional verifiers.
-
GovernmentBusiness Private-sector data shared with development organizations (World Bank, IMF, IDB).
-
Aggregated lists Version-controlled databases — including the hospital-price transparency corpus.
-
AcademicArchives
-
EnergyAPIs Energy think tank Ember's open electricity data: yearly generation, demand, capacity, and power-sector emissions for 200+ countries and monthly data for about 90, through CSV downloads, a data explorer, and a free-key API under CC BY 4.0.
-
Benchmarks & evalsScience CC BY data hubs tracking the AI frontier - a 3,600+ model database with training compute, a Benchmarking Hub (GPQA Diamond, SWE-bench Verified, its own FrontierMath), and infrastructure hubs for AI data centers, GPU clusters, AI chip sales, and ML hardware - behind most 'AI progress' charts.
-
Responsible AIGovernment Georgetown CSET's AI Governance and Regulatory Archive: 1,000+ AI laws, regulations, and standards with summaries, full text, and tags for risks, harms, and governance strategies. The bulk dataset is refreshed monthly on Zenodo under CC BY-NC 4.0.
-
APIs Web search built as an API for AI applications.
-
Benchmarks & evalsAcademicSecurity & privacy Carnegie Mellon's capability-ladder benchmark for AI security agents: 41 patched Chromium V8 bugs graded on 16 deterministic milestones, from reaching vulnerable code to code execution, so results show how far a model climbs rather than pass/fail.
-
Synthetic data The standard library for generating fake records in tests and demos.
-
Benchmarks & evals AutoGluon's realistic forecasting benchmark: 100 tasks across seven domains, 46 with covariates, ranking 30+ models by win rate and skill score with bootstrapped confidence intervals; the fev library reproduces every task.
-
AcademicScience
-
Responsible AIBenchmarks & evalsAcademic Stanford CRFM's scorecard of foundation model developers on 100 transparency indicators covering training data, compute, usage policies, and downstream impact. The December 2025 edition scored 13 companies, with indicator-level data on GitHub.
-
APIsGeospatial
-
Social scienceArchives Global news events, tone, and knowledge graphs back to 1979; free raw files and BigQuery tables.
-
Security & privacyGovernment Law firm CMS's database of publicly known GDPR fines, including UK GDPR and some ePrivacy penalties: 3,200+ actions filterable by country, sector, GDPR article, and amount, with summary statistics. Free.
-
Benchmarks & evalsAcademic Human-aligned image-editing benchmark from Nanyang Technological University, StepFun, and Southeast University: 1,200 real user edit requests across 23 tasks, judged for instruction following, visual quality, and consistency with the source image, and ranked by Elo.
-
Benchmarks & evals Salesforce's general time-series forecasting benchmark: 23 datasets, 144k series, and 177M points across seven domains and ten frequencies, with a public leaderboard and a non-leaking pretraining corpus.
-
EnergyGeospatial Global Energy Monitor's 25 trackers of energy infrastructure, including a power tracker covering 180,000+ facilities in 200 countries and areas with capacity, status, owner, fuel, and location. Tracker data is CC BY 4.0.
-
Climate & environment
-
Responsible AIGovernmentSocial science The Global Center on AI Governance's country ranking: 135 countries scored on 38 responsible-AI indicators across five dimensions, including inclusion and diversity, built from 68,000+ data points, with data downloads and an evidence explorer. Second edition July 2026.
-
BusinessAggregated lists
-
APIsAggregated lists
-
GeospatialClimate & environment Petabytes of analysis-ready satellite and environmental data.
-
APIsGeospatial
-
ArchivesScience
-
AcademicScience Open repository for research data — 100k+ datasets across every field.
-
Security & privacyAPIs Troy Hunt's index of data breaches: 1,000+ breached sites and roughly 17.8 billion exposed accounts. Breach metadata is free through the API under CC BY 4.0; email and domain searches need a paid key, and Pwned Passwords stays free.
-
Synthetic data
-
Benchmarks & evalsAcademicResponsible AI Holistic Evaluation of Language Models: transparent multi-metric leaderboards with every prompt and raw prediction published, including HELM Safety (harmful requests, bias, and over-refusal) and AIR-Bench (risk categories drawn from regulations and company policies).
-
GovernmentSocial science UN OCHA’s portal — crisis and development data for every country.
-
Benchmarks & evalsAcademic CAIS and Scale AI's 2,500-question exam at the frontier of human knowledge, written by nearly 1,000 subject experts across 100+ subjects; results table on the site, dataset on Hugging Face (cais/hle).
-
GovernmentAggregated lists Federated search across thousands of Opendatasoft-hosted portals. Opendatasoft rebranded to Huwise in 2025; same federated hub.
-
Aggregated lists
-
Health
-
Benchmarks & evalsAcademic The corpus that started the deep-learning era: 14M labeled images and the ILSVRC challenge lineage.
-
Entertainment
-
Transit Commercial traffic and mobility data.
-
CensusSocial science Harmonized census and survey microdata — US and international, decades deep.
-
APIs Turns any URL into LLM-ready markdown.
-
Aggregated lists
-
Science Explore label mistakes in the ten most-cited ML benchmarks.
-
ScienceAggregated lists Funds and catalogs ML datasets for underserved languages and regions.
-
APIsEntertainment
-
Benchmarks & evalsAcademic Contamination-limited LLM benchmark from Abacus.AI, NYU, Nvidia, Maryland, and USC: monthly-refreshed questions across math, coding, reasoning, data analysis, instruction following, and language, scored against objective ground truth without an LLM judge.
-
Benchmarks & evals Crowd-sourced human-preference leaderboard for AI models (née Chatbot Arena, now at arena.ai): millions of blind head-to-head votes across text, text-to-image, image editing, text-to-video, image-to-video, and video editing, with conversation datasets released for research.
-
Benchmarks & evalsAcademic The research org behind Chatbot Arena, Vicuna, and SGLang; publishes open conversation datasets like LMSYS-Chat-1M.
-
Aggregated lists The linked-open-data cloud diagram and registry.
-
Responsible AIAcademic MIT's living database of 1,700+ AI risks extracted from 65 frameworks and coded by cause and domain, shared as a copyable spreadsheet under CC BY 4.0, with a companion tracker that classifies AI Incident Database incidents.
-
Benchmarks & evalsScienceResponsible AI Home of the MLPerf suites - training, inference, and storage results - plus open datasets like People's Speech. AILuminate grades AI chat systems on safety across hazard categories, and MLPerf Inference v6.0 (2026) added a text-to-video test that times generation and uses VBench as the quality check.
-
Science Benchmark time-series collections for forecasting research.
-
Synthetic data
-
Science Community datasets — Common Voice and beyond.
-
Benchmarks & evalsAcademic Carnegie Mellon's live arena for text-to-music: listeners vote blind between two generated tracks for the same prompt, ranked on separate instrumental and vocal boards, with the votes, prompts, and audio released monthly on Hugging Face under CC BY 4.0. Not the same project as Artificial Analysis's Music Arena.
-
HealthScience
-
ScienceAggregated lists Interactive repository of thousands of graph and network datasets.
-
APIs
-
Responsible AIGovernment OECD.AI's monitor of about 17,000 AI incidents and hazards detected in global news coverage and classified by harm type, severity, affected stakeholders, and country, with downloadable filtered results. A beta, and not an official OECD position.
-
GovernmentBusiness
-
Benchmarks & evalsAcademic Shanghai Jiao Tong University and StepFun's text-to-image benchmark: prompt-image alignment, text rendering, reasoning, style, and diversity in English and Chinese, with a per-metric leaderboard.
-
Aggregated listsGovernment A map of 2,600+ open-data portals worldwide.
-
Aggregated lists Q&A site where dataset hunts get answered.
-
Benchmarks & evalsAPIs Revealed preference rather than a test set: live token-share rankings of AI models from real OpenRouter traffic.
-
Geospatial The open map of the world: planet dumps, regional extracts, and the ODbL data behind countless products.
-
Benchmarks & evalsAcademic XLANG Lab's computer-use benchmark: agents operate real Ubuntu, Windows, and macOS desktops; OSWorld-Verified is the audited leaderboard and OSWorld 2.0 adds 108 long-horizon professional tasks.
-
Benchmarks & evalsScience Google DeepMind and INSAIT's test of whether video generation models understand physics: models continue real filmed experiments in solid mechanics, fluid dynamics, optics, thermodynamics, and magnetism, and are scored against what actually happened. Leaderboard on the site, code and data on GitHub.
-
APIsAggregated lists
-
APIsAggregated lists
-
Security & privacyAPIs Independent tracker of victims named on ransomware groups' leak sites since 2022: 30,000+ victims across nearly 400 groups, with a free API, JSON data, and RSS. It hosts no leaked data, and commercial reuse needs permission.
-
APIs
-
AcademicAggregated lists Registry of 3,000+ research data repositories — find the repository for any field.
-
Benchmarks & evalsAggregated lists Scale AI's expert-built evaluation hub: 20+ leaderboards spanning agentic coding (SWE-Bench Pro, SWE Atlas, MCP Atlas), frontier reasoning (Humanity's Last Exam), and safety, mixing private held-out sets with open ones.
-
Science Peer-reviewed descriptors of high-value datasets.
-
Benchmarks & evalsAcademic Coding agents graded like senior engineers: under-specified feature and bug tasks that need runtime investigation, scored on code quality as well as correctness (Princeton, UW-Madison, and Snorkel AI; 50 public + 50 private tasks).
-
APIs
-
GeospatialCensus
-
APIsGovernment
-
Benchmarks & evalsAcademic Song aesthetics dataset from Northwestern Polytechnical University's ASLP lab: 2,399 full English and Chinese songs (about 140 hours) rated by 16 musically trained annotators for coherence, memorability, vocal phrasing, structure, and musicality, with an open scoring toolkit that song-generation papers report against. CC BY-NC-SA 4.0 (non-commercial).
-
Archives
-
Archives Run SQL against Stack Overflow live, in the browser.
-
ScienceSocial scienceResponsible AI Stanford HAI's annual AI Index. The data behind its chapters is public, including research, the economy, policy, public opinion, and responsible AI (safety, fairness, transparency).
-
BusinessAggregated lists
-
Benchmarks & evals Can models fix real GitHub issues? The coding-agent benchmark of record, with Verified and Multimodal leaderboards.
-
Synthetic data
-
Synthetic data R package for synthetic versions of sensitive microdata.
-
Benchmarks & evalsAcademic Text-to-image benchmark from USTC, Kuaishou's Kling team, and the University of Hong Kong that separates composition (objects, attributes, relations, text rendering) from reasoning (deductive, inductive, abductive): 1,080 prompts and about 13,500 checklist questions, with a leaderboard of open and closed models.
-
Benchmarks & evals Evaluates AI agents on end-to-end tasks in a real terminal, from compiling code to training models.
-
Benchmarks & evalsScience Terminal-Bench's sibling for research workflows: AI agents on end-to-end scientific tasks across domains, from Stanford, Harbor, and the Laude Institute, with a public leaderboard and tasks on Harbor Hub.
-
HealthCensus Demographic and Health Surveys — household survey microdata for 90+ countries.
-
APIsSports
-
Benchmarks & evalsAggregated lists Independent daily ranking of 34k+ open-source agent skills from GitHub - stars, commits, and SKILL.md analysis.
-
APIsEntertainment
-
Benchmarks & evalsAcademic Zero-shot time-series foundation-model benchmark built on 50 fresh datasets and 98 forecasting tasks kept out of pretraining corpora, with a multi-granular leaderboard (one of the three suites Google used to rank TimesFM-3).
-
Science Twice-yearly ranking of the world's 500 most powerful supercomputers since 1993, each list downloadable, with a sublist generator by country, vendor, processor, and interconnect and the companion Green500 energy-efficiency ranking.
-
TransitGeospatial Aggregates GTFS feeds from thousands of transit operators.
-
GovernmentSocial science
-
Benchmarks & evals Community-run blind arena for text-to-speech on Hugging Face: type a line, hear two anonymous models read it back, and pick the one that sounds more human; the votes build an open leaderboard.
-
HealthScience
-
GovernmentCensus
-
GovernmentSocial science
-
Benchmarks & evalsAcademic The UnifiedReward team's prompt-following benchmark for text-to-image: 600 prompts per language in English and Chinese, short and long forms, checked against 10 primary and 27 sub-dimensions by Gemini 2.5 Pro or an open evaluation model. This Space is the English leaderboard.
-
Benchmarks & evalsAcademic The most-cited video-generation benchmark, from Nanyang Technological University and Shanghai AI Lab's Vchitect team: text-to-video and image-to-video models scored on 16 quality dimensions, with VBench-2.0 adding physics, commonsense, and human fidelity. MLPerf's text-to-video test uses it as the quality check.
-
Benchmarks & evalsAcademic Meituan LongCat and Fudan University's multi-turn benchmark for interactive video world models: 289 test cases and 22 metrics covering video quality, adherence to the scene setup, response to actions, consistency, and physical plausibility, with a leaderboard of commercial and open models.
-
Benchmarks & evalsClimate & environment Google Research's benchmark for data-driven global weather models: headline scorecards for 1-15 day forecasts (RMSE, CRPS, SEEPS precipitation) plus cloud-optimized ERA5 ground truth and baseline datasets.
-
APIsBusiness Crawled web data as a feed — news, blogs, forums, dark web.
-
Health
-
Archives
-
Aggregated lists
-
Benchmarks & evalsAcademic Song-generation benchmark from the Multimodal Art Projection (M-A-P) community: 192 English and Chinese prompts run through 17 systems including Suno and Mureka, scored by automatic metrics for song quality, prompt control, lyric accuracy, and text-audio match. Prompts, results, and seeds are published; audio is not. Released in September 2026 alongside M-A-P's own YuE2 model, which ranks first, and no license is stated yet.
-
BusinessGovernment
-
Benchmarks & evalsAcademic Stanford's unified benchmark for world generation: 3D, 4D, image-to-video, and text-to-video models build scenes along specified camera paths across 3,000 test cases, scored for controllability, quality, and dynamics. Leaderboard and dataset on Hugging Face.