Talk:Common Crawl
Add topic| This article is rated Start-class on Wikipedia's content assessment scale. It is of interest to the following WikiProjects: | ||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||
| The Wikimedia Foundation's Terms of Use require that editors disclose their "employer, client, and affiliation" with respect to any paid contribution; see WP:PAID. For advice about reviewing paid contributions, see WP:COIRESPONSE. Edits made by the below user(s) were last checked for neutrality on 10 May 2026 by Superb Owl.
|
External links modified
[edit]Hello fellow Wikipedians,
I have just modified one external link on Common Crawl. Please take a moment to review my edit. If you have any questions, or need the bot to ignore the links, or the page altogether, please visit this simple FaQ for additional information. I made the following changes:
- Added archive https://web.archive.org/web/20150404132256/http://blog.commoncrawl.org/ to http://blog.commoncrawl.org/
When you have finished reviewing my changes, you may follow the instructions on the template below to fix any issues with the URLs.
This message was posted before February 2018. After February 2018, "External links modified" talk page sections are no longer generated or monitored by InternetArchiveBot. No special action is required regarding these talk page notices, other than regular verification using the archive tool instructions below. Editors have permission to delete these "External links modified" talk page sections if they want to de-clutter talk pages, but see the RfC before doing mass systematic removals. This message is updated dynamically through the template {{source check}} (last update: 5 June 2024).
- If you have discovered URLs which were erroneously considered dead by the bot, you can report them with this tool.
- If you found an error with any archives or the URLs themselves, you can fix them with this tool.
Cheers.—InternetArchiveBot (Report bug) 10:09, 11 August 2017 (UTC)
Some statistics by CommonCrawl
[edit]https://commoncrawl.github.io/cc-crawl-statistics/ Should these percentages be added? — Preceding unsigned comment added by 5.206.101.180 (talk) 18:30, 18 May 2022 (UTC)
There is another important page of statistics
[edit]Here it is: https://commoncrawl.github.io/cc-crawl-statistics/plots/charsets.html 5.206.107.6 (talk) 16:18, 19 May 2023 (UTC)
Size discrepancy?
[edit]The lead says the data set is several petabytes but the last record in the timeline says it's 386 TB. Which is it? 179.24.10.81 (talk) 14:46, 23 October 2024 (UTC)
- There is no discrepancy: April 2024 crawl is 386 TB, while the sum of all crawls is several PB. MGeog2022 (talk) 12:05, 19 April 2025 (UTC)
- Now over 10 petabytes. These numbers are all published on our website. Greg (talk) 03:49, 15 February 2026 (UTC)
Disagreements about CCF
[edit]So I'm a very early contributor to Wikipedia, but I'm a bit rusty with my skills: is there something I should do when a journalist accuses me of lying? I added my non-profit's response to the false accusation. But is there something else I should do? Greg (talk) 05:12, 15 February 2026 (UTC)
- Hi @Greg Lindahl — thanks for your contributions (it's cool to see someone who joined in 2001 still active!) and for reaching out. Because you work for Common Crawl, you have a paid conflict of interest and should follow the guidance for paid editors at Wikipedia:Conflict of interest. This includes, importantly, not directly editing the article, and instead proposing changes here on the talk page. You can use the {{Edit request}} template to request that an independent editor review them.
- I have procedurally reverted the article to the state it was in before you edited it. This is not a direct response to the content of the edits themselves, but rather a process step to ensure that the changes you are seeking are reviewed by independent editors. For the page move, your request will be more likely to be accepted if you demonstrate that "Common Crawl Foundation" has become the common name. And for the response to the Atlantic, your request will be more likely to be accepted if you show that the response has been covered in reliable sources rather than just posted to your website (a primary source).
- Cheers, Sdkb talk 13:46, 14 April 2026 (UTC)
- Thanks. I have no idea how this works,and I'm happy to have this edit dropped on the floor, if you prefer that. Greg (talk) 23:44, 9 May 2026 (UTC)
- In terms of the common name debate, Common Crawl Foundation is used here: The Register 2025, LA Times 2012, The Atlantic 2025 (mostly referred to as Common Crawl), Washington Post 2025. Superb Owl (talk) 18:38, 10 May 2026 (UTC)
- Thanks. I have no idea how this works,and I'm happy to have this edit dropped on the floor, if you prefer that. Greg (talk) 23:44, 9 May 2026 (UTC)
Proposed additions to "Refined versions"
[edit]| The user below has a request that an edit be made to Common Crawl. That user has an actual or apparent conflict of interest. The requested edits backlog is very high. Please be extremely patient. There are currently 883 requests waiting for review. Please read the instructions for the parameters used by this template for accepting and declining them, and review the request below and make the edit if it is well sourced, neutral, and follows other Wikipedia guidelines and policies. Remember to set the |answered= parameter to "yes" when the request has been accepted, rejected or on hold awaiting user input. |
Disclosure: I am employed by amber Tech GmbH. One of the two authors of the GC4 corpus in part 3 below is a co-founder of that company. I can see from this page that earlier paid edits to the article were procedurally reverted, so I am bringing everything here rather than editing the article myself — including parts 1 and 2, which have no connection to my employer. Please take or leave any of it as you see fit.
Part 1 — expand the first sentence of "Refined versions"
The section currently names only FineWeb, DCLM and C4. Proposed replacement for the existing sentence:
A number of organizations take raw Common Crawl data and refine it into datasets that exclude edgy content or are otherwise higher-quality for their purposes, such as FineWeb, DCLM and C4.<ref name=":1" /> Other examples include CCNet, a pipeline published in 2019 that deduplicates Common Crawl documents, identifies their language and filters them by similarity to Wikipedia text,<ref name="wenzek2020" /> and OSCAR, a multilingual corpus produced by filtering and classifying Common Crawl snapshots by language.<ref name="ortizsuarez2019" /> mC4, a multilingual counterpart to C4 covering 101 languages, was built for training the mT5 model.<ref name="xue2021" />
Part 2 — one sentence on the C4 documentation study
To be placed in "Colossal Clean Crawled Corpus", after the sentence beginning "As of 2023, there were some concerns…" and before the sentence on the 2024 study. Dodge et al. is the peer-reviewed basis for much of what the existing journalistic sources there report:
A 2021 documentation study examined C4, which was built from a single Common Crawl snapshot, and found substantial text from unexpected sources such as patents and US military websites as well as machine-generated text. The same study showed that the blocklist filtering used to create the corpus disproportionately removed text from and about minority individuals.<ref name="dodge2021" />
Part 3 — the part I have a conflict of interest about
The section currently covers English-language derivatives only. GC4 is, as far as I know, the largest freely available German one, which seems to me a genuine gap — but I am not a neutral party here, so please judge it on the merits or simply decline. To be appended after the mC4 sentence:
Language-specific derivatives include GC4 (German colossal, cleaned Common Crawl corpus), released in 2021 by Philipp Reissel and Philip May, which Bornheim et al. describe as comprising about 540 GB of German text.<ref name="bornheim2021" />
The sentence is deliberately limited to what an independent paper states; the corpus authors' own project page is not used as a source. If you include parts 1 and 2 and leave out part 3, that is entirely fine with me — I will not reinstate it.
Sources
- CCNet: Wenzek, Lachaux, Conneau, Chaudhary, Guzmán, Joulin, Grave: CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data. LREC 2020. arXiv:1911.00359
- OSCAR: Ortiz Suárez, Sagot, Romary: Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures. CMLC-7, 2019, pp. 9–16
- mC4: Xue et al.: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer. NAACL-HLT 2021. arXiv:2010.11934
- Dodge et al.: Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. EMNLP 2021. arXiv:2104.08758
- GC4: Bornheim, Grieger, Bialonski: FHAC at GermEval 2021. arXiv:2109.02966
I can supply full citation templates for any of these if that is useful. AI-Nerd-42 (talk) 07:02, 19 September 2026 (UTC)
- Start-Class organization articles
- Mid-importance organization articles
- WikiProject Organizations articles
- Start-Class Artificial Intelligence articles
- Mid-importance Artificial Intelligence articles
- WikiProject Artificial Intelligence articles
- Start-Class Internet articles
- Mid-importance Internet articles
- WikiProject Internet articles
- Start-Class Digital Preservation articles
- Talk pages of subject pages with paid contributions
- Wikipedia conflict of interest edit requests