Stackmark › Datasets

People's Speech

Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech.

  • Type: Dataset
  • Imported from Hugging Face
  • Popularity: 286 likes on Hugging Face
  • License: cc-by-2.0
  • Source: https://huggingface.co/datasets/MLCommons/peoples_speech
  • Tags: annotations_creators:crowdsourced, annotations_creators:machine-generated, language_creators:crowdsourced, language_creators:machine-generated, multilinguality:monolingual, source_datasets:original, robust-speech-recognition, noisy-speech-recognition, speech-recognition
  • Updated: 2026-10-01

More datasets

  • prompts.chat — a.k.a. Awesome ChatGPT Prompts This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts. 📢 Notice This Huggi…
  • FineWeb — 🍷 FineWeb 15 trillion tokens of the finest data the 🌐 web has to offer What is it? The 🍷 FineWeb dataset consists of more than 18.5T tok…
  • hh-rlhf — Dataset Card for HH-RLHF Dataset Summary This repository provides access to two different kinds of data: Human preference data about helpfu…
  • Grade School Math 8K — Dataset Card for GSM8K Dataset Summary GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school m…
  • wikipedia — Dataset Card for Wikimedia Wikipedia Dataset Summary Wikipedia dataset containing cleaned articles of all languages. The dataset is built f…
  • OpenAssistant Conversations — OpenAssistant Conversations Dataset (OASST1) Dataset Summary In an effort to democratize research on large-scale alignment, we release Open…
  • Alpaca — Dataset Card for Alpaca Dataset Summary Alpaca is a dataset of 52,000 instructions and demonstrations generated by OpenAI's text-davinci-00…
  • databricks-dolly-15k — Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in s…

Privacy · Terms · llms.txt