Corpus Info

Available corpora

Corpora at NWU have mostly been developed in collaboration with various corpus suppliers, such as publishing houses, news websites, (literary) blog sites, libraries, etc. These corpora may be used for academic research purposes only. For commercial licencing options, please contact us directly.

These corpora are divided into two collections:

Available interfaces / platforms

COCO@NWU hosts the corpora on two platforms:

COCO@NWU Regular

Size
  • Total words: 370 179 957
  • Total tokens: 426 555 786
  • Frequency: "die": 27 342 482

A collection of government documents from the various websites and electronic publications of the South African government.

Size
  • Total words: 1 514 385
  • Total tokens: 1 736 282
  • Frequency: "die": 128 461
Reference

Department of Sports, Arts and Culture & CTexT. 2026. NCHLT- Afrikaanse Korpus 2.0 [NCHLT Afrikaans Corpus 2.0]. Potchefstroom: CTexT, North-West University. ISLRN: 544-932-849-161-3. Available at: https://hdl.handle.net/20.500.12185/293

A subset of the Leipzig Corpora Collection.

Subset contains Afrikaans texts up until 2023. © 2025 Abteilung Automatische Sprachverarbeitung, Universität Leipzig.

Size
  • Total words: 137 094 849
  • Total tokens: 156 170 924
  • Frequency: "die": 9 442 681
Reference

Goldhahn, D. Eckart, T. & Quasthoff, U. 2012. Building Large Monolingual Dictionaries at the Leipzig Corpora Collection: From 100 to 200 Languages. In: Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12).

A collection of original plays submitted for the ATKV Youth Theatre Competition.

Collection covers competition entries from 2016 to 2022.

Size
  • Total words: 4 007 970
  • Total tokens: 5 025 644
  • Frequency: "die": 155 470
Reference

ATKV & CTexT. 2026. NWU/ATKV-Tienertoneelkorpus 2.0 [NWU/ATKV Theatre for Teenagers Corpus 2.0]. Potchefstroom: CTexT, North-West University.

A collection of news articles and blogs as published on the media house Maroela Media's website.

Collection covers all material up to the end of January 2026.

Size
  • Total words: 58 266 623
  • Total tokens: 67 104 583
  • Frequency: "die": 4 239 412
Reference

Maroela Media & CTexT. 2026. NWU/Maroela Mediakorpus 3.0 [NWU/Maroela Media Corpus 3.0]. Potchefstroom: CTexT, North-West University.

A corpus of Afrikaans books (mostly fiction) published by Lapa Publishers.

Size
  • Total words: 19 981 313
  • Total tokens: 23 607 580
  • Frequency: "die": 908 282
Reference

LAPA Uitgewers & CTexT. 2026. NWU/LAPA-Korpus 2.0 [NWU/LAPA Corpus 2.0]. Potchefstroom: CTexT, North-West University.

A collection of articles published in the Taalgenoot magazine.

Collection covers all material from 2006 to 2025.

Size
  • Total words: 2 925 479
  • Total tokens: 3 422 587
  • Frequency: "die": 171 574
Reference

ATKV & CTexT. 2026. NWU/ATKV-Taalgenootkorpus 2.0 [NWU/ATKV Taalgenoot Corpus 2.0]. Potchefstroom: CTexT, North-West University.

A corpus of Afrikaans books (fiction and non-fiction) published by the publisher Protea Boekhuis.

Size
  • Total words: 14 093 645
  • Total tokens: 16 395 760
  • Frequency: "die": 959 753
Reference

Protea Boekhuis & CTexT. 2026. NWU/Protea Boekhuiskorpus 3.0 [NWU/Protea Boekhuis Corpus 3.0]. Potchefstroom: CTexT, North-West University.

A collection of news bulletins, as broadcast on Radio Sonder Grense and published on their website.

Collection covers all material from January 2005 to January 2026.

Size
  • Total words: 49 052 380
  • Total tokens: 54 073 952
  • Frequency: "die": 4 849 459
Reference

Radio Sonder Grense & CTexT. 2026. NWU/RSG-nuuskorpus 3.0 [NWU/RSG News Corpus 3.0]. Potchefstroom: CTexT, North-West University.

A stratified corpus used by the Afrikaans Language Commission, consisting of a variety of genres and domains, including academic journals, newspapers, literary works, informal writings, etc.

Size
  • Total words: 46 470 173
  • Total tokens: 53 643 030
  • Frequency: "die": 3 623 055
Reference

Taalkommissie van die Suid-Afrikaanse Akademie vir Wetenskap en Kuns. 2011. Taalkommissiekorpus 2.0 [Language Commission Corpus 2.0]. Potchefstroom: CTexT, North-West University.

A collection of documents from the various web pages available on the Afrikaans version of Wikipedia, Wikibooks, Wikiquote and Wiktionary.

Collection covers all material up to August 2025.

Size
  • Total words: 30 748 437
  • Total tokens: 38 471 311
  • Frequency: "die": 2 633 538
Reference

Wikimedia Foundation Inc. & CTexT. 2026. NWU/Wikimedia Afrikaanse korpus 2.0 [NWU/Wikimedia Afrikaans Corpus 2.0]. Potchefstroom: CTexT, North-West University.

The data included in the corpora presented here issue from Pula/Imvula, a South African magazine focusing on the developing farmer and published by Grain SA (https://www.grainsa.co.za/farmer-development). The main aim of the magazine is to support developing farmers in becoming sustainable commercial farmers. The magazine is distributed on a monthly basis and currently only available in English. However, from 2007 until September 2024 it was published in five languages (English, isiXhosa, isiZulu, Sesotho, and Setswana), and previously (2007-2019) also included Afrikaans and Sesotho sa Leboa (discontinued due to lack of funding).

Size
  • Total words: 706 337
  • Total tokens: 798 067
  • Frequency: "die": 48 560
Reference

Pula Imvula & CTexT. 2025. NWU/Pula Imvula Afrikaans Corpus 1.0. Potchefstroom: CTexT, North-West University.

Klyntji.com is an independent, online journal published in all variants of Afrikaans, focusing on diverse and progressive art and culture from the global South. The corpus covers material from 2014 to 2026.

Size
  • Total words: 821 203
  • Total tokens: 945 451
  • Frequency: "die": 43 805
Reference

Klyntji & CTexT. 2026. NWU/Klyntji Corpus 1.1. Potchefstroom: CTexT, North-West University.

Vrye Weekblad was a progressive Afrikaans national weekly newspaper published from 1988 to 1994. It was relaunched as an online newspaper from 6 April 2019 to 28 March 2025. The corpus includes material from 2019 to 2025, but excludes user comments.

Size
  • Total words: 4 497 163
  • Total tokens: 5 160 615
  • Frequency: "die": 138 432
Reference

Vrye Weekblad & CTexT. 2025. NWU/Vrye Weekblad-korpus 1.0 [NWU/Vryeweekblad Corpus 1.0]. Potchefstroom: CTexT, North-West University.

Size
  • Total words: 1 028 794
  • Total tokens: 1 166 682
  • Frequency: "the": 65 232

The data included in the corpora presented here issue from Pula/Imvula, a South African magazine focusing on the developing farmer and published by Grain SA (https://www.grainsa.co.za/farmer-development). The main aim of the magazine is to support developing farmers in becoming sustainable commercial farmers. The magazine is distributed on a monthly basis and currently only available in English. However, from 2007 until September 2024 it was published in five languages (English, isiXhosa, isiZulu, Sesotho, and Setswana), and previously (2007-2019) also included Afrikaans and Sesotho sa Leboa (discontinued due to lack of funding).

Size
  • Total words: 1 028 794
  • Total tokens: 1 166 682
  • Frequency: "the": 65 232
Reference

Pula Imvula & CTexT. 2025. NWU/Pula Imvula English Corpus 1.0. Potchefstroom: CTexT, North-West University.

Size
  • Total words: 904 108
  • Total tokens: 1 014 905
  • Frequency: "go": 52 179

The data included in the corpora presented here issue from Pula/Imvula, a South African magazine focusing on the developing farmer and published by Grain SA (https://www.grainsa.co.za/farmer-development). The main aim of the magazine is to support developing farmers in becoming sustainable commercial farmers. The magazine is distributed on a monthly basis and currently only available in English. However, from 2007 until September 2024 it was published in five languages (English, isiXhosa, isiZulu, Sesotho, and Setswana), and previously (2007-2019) also included Afrikaans and Sesotho sa Leboa (discontinued due to lack of funding).

Size
  • Total words: 904 108
  • Total tokens: 1 014 905
  • Frequency: "go": 52 179
Reference

Pula Imvula & CTexT. 2025. NWU/Pula Imvula Sesotho sa Leboa Corpus 1.0. Potchefstroom: CTexT, North-West University.

Size
  • Total words: 1 235 476
  • Total tokens: 1 376 803
  • Frequency: "ho": 71 856

The data included in the corpora presented here issue from Pula/Imvula, a South African magazine focusing on the developing farmer and published by Grain SA (https://www.grainsa.co.za/farmer-development). The main aim of the magazine is to support developing farmers in becoming sustainable commercial farmers. The magazine is distributed on a monthly basis and currently only available in English. However, from 2007 until September 2024 it was published in five languages (English, isiXhosa, isiZulu, Sesotho, and Setswana), and previously (2007-2019) also included Afrikaans and Sesotho sa Leboa (discontinued due to lack of funding).

Size
  • Total words: 1 235 476
  • Total tokens: 1 376 803
  • Frequency: "ho": 71 856
Reference

Pula Imvula & CTexT. 2025. NWU/Pula Imvula Sesotho Corpus 1.0. Potchefstroom: CTexT, North-West University.

Size
  • Total words: 1 218 908
  • Total tokens: 1 335 498
  • Frequency: "go": 95 722

The data included in the corpora presented here issue from Pula/Imvula, a South African magazine focusing on the developing farmer and published by Grain SA (https://www.grainsa.co.za/farmer-development). The main aim of the magazine is to support developing farmers in becoming sustainable commercial farmers. The magazine is distributed on a monthly basis and currently only available in English. However, from 2007 until September 2024 it was published in five languages (English, isiXhosa, isiZulu, Sesotho, and Setswana), and previously (2007-2019) also included Afrikaans and Sesotho sa Leboa (discontinued due to lack of funding).

Size
  • Total words: 1 218 908
  • Total tokens: 1 335 498
  • Frequency: "go": 95 722
Reference

Pula Imvula & CTexT. 2025. NWU/Pula Imvula Setswana Corpus 1.0. Potchefstroom: CTexT, North-West University.

Size
  • Total words: 712 468
  • Total tokens: 825 364
  • Frequency: "ukuba": 10 398

The data included in the corpora presented here issue from Pula/Imvula, a South African magazine focusing on the developing farmer and published by Grain SA (https://www.grainsa.co.za/farmer-development). The main aim of the magazine is to support developing farmers in becoming sustainable commercial farmers. The magazine is distributed on a monthly basis and currently only available in English. However, from 2007 until September 2024 it was published in five languages (English, isiXhosa, isiZulu, Sesotho, and Setswana), and previously (2007-2019) also included Afrikaans and Sesotho sa Leboa (discontinued due to lack of funding).

Size
  • Total words: 712 468
  • Total tokens: 825 364
  • Frequency: "ukuba": 10 398
Reference

Pula Imvula & CTexT. 2025. NWU/Pula Imvula isiXhosa Corpus 1.0. Potchefstroom: CTexT, North-West University.

Size
  • Total words: 661 442
  • Total tokens: 774 127
  • Frequency: "futhi": 8 489

The data included in the corpora presented here issue from Pula/Imvula, a South African magazine focusing on the developing farmer and published by Grain SA (https://www.grainsa.co.za/farmer-development). The main aim of the magazine is to support developing farmers in becoming sustainable commercial farmers. The magazine is distributed on a monthly basis and currently only available in English. However, from 2007 until September 2024 it was published in five languages (English, isiXhosa, isiZulu, Sesotho, and Setswana), and previously (2007-2019) also included Afrikaans and Sesotho sa Leboa (discontinued due to lack of funding).

Size
  • Total words: 661 442
  • Total tokens: 774 127
  • Frequency: "futhi": 8 489
Reference

Pula Imvula & CTexT. 2025. NWU/Pula Imvula isiZulu Corpus 1.0. Potchefstroom: CTexT, North-West University.

COCO@NWU Special

Size
  • Total words: 72 235 291
  • Total tokens: 83 027 579
  • Frequency: "die": 2 795 800

Collection of unedited user comments on web pages, arranged in subcorpora per web page.

Current collection: Comments from two different websites (anonymous).

Size
  • Total words: 51 044 243
  • Total tokens: 58 262 184
  • Frequency: "die": 2 639 445
Reference

CTexT. 2026. NWU Kommentaarkorpus 3.0 [NWU Comment Corpus 3.0]. Potchefstroom: CTexT, North-West University.

Collection of historical Afrikaans documents, arranged by subcorpus.

Current collection: VOC documents from 1662 to 1791, as digitized by the Tracing History Trust.

Size
  • Total words: 19 343 422
  • Total tokens: 22 647 357
  • Frequency: "die": 67 177
Reference

Tracing History Trust & CTexT. 2021 . NWU/THT korpus 1.4 [NWU/THT Corpus 1.4]. Potchefstroom: CTexT, North-West University.

Collection of informal blogs as published on watkykjy.co.za up to and including the end of February 2025.

Size
  • Total words: 1 847 626
  • Total tokens: 2 118 038
  • Frequency: "die": 89 178
Reference

WatKykJy & CTexT. 2026. NWU/WatKykJy-korpus 3.0 [NWU/WatKykJy Corpus 3.0]. Potchefstroom: CTexT, North-West University.

Automatic corpus processing

For both search platforms, all non-English corpora have been automatically processed, annotated, and (re)compiled using in-house tools developed by CTexT. NOTE: No deduplication has been done on the Leipzig corpus.

Processing and annotation of Afrikaans texts includes the following levels:

Afrikaans Sentence separation and tokenisation

For more information on the algorithms and resources used, please cite:

Puttkammer, M.J. 2006. Outomatiese Afrikaanse tekseenheididentifisering [Automatic Afrikaans tokenisation, Sentence separation and named-entity recognition]. MA thesis, North-West University.

Note that Sentence separation and tokenisation counts might differ slightly between Afrikaans corpora that users have uploaded themselves on Sketch Engine, and those that have been uploaded by NWU on the Sketch Engine and COCO platforms. The main reason for this is that Sketch Engine and our in-house tools use slightly different definitions of what sentences and words are.

Afrikaans Part of speech (POS) tagging

  • Accuracy: 95,7%
  • Tagset
  • Please cite: Eiselen, R. & Puttkammer, M.J. 2014. Developing Text Resources for Ten South African Languages. In: Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'14) (pp. 3698-3703).

Afrikaans Named-entity recognition (NER)

  • F-score: 0,76
  • Tagset: See below
  • Please cite: Eiselen, R. 2016. Government domain named entity recognition for South African languages. In: Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16) (pp. 3344-3348).

Afrikaans Phrase chunking (PC)

  • F-score: 0,95
  • Tagset: See below
  • Please cite: Eiselen, R. 2016. South African language resources: phrase chunking. In: Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16) (pp. 689-693).

Afrikaans POS tagset

Accessed at: https://hlt.nwu.ac.za/docs/help/pos_tags/AF/AFpos.htm

Afrikaans NER tagset

Tag Meaning
B-/I-ORG Organisation
B-/I-PER Person
B-/I-LOC Location
B-/I-MISC Miscellaneous
OUT Outside

Afrikaans PC tagset

Tag Meaning
B-/I-NOUN Noun phrase
B-/I-VERB Verb phrase
B-/I-PREP Prepositional phrase
B-/I-ADJ Adjective phrase
B-/I-ADV Adverbial phrase
OUT Outside

Citations and references

Initiative/website

COCO@NWU. 2025. Corpus Cooperative at North-West University. Available at: http://coco.nwu.ac.za.

Corpus collections

Choose between one or more of the following:

Annotation / processing tools

See the citation information under each of the tools mentioned above.

Sample texts

You may use / adapt the texts below in your research proposal, ethics application, article, thesis, etc. They merely serve as examples how you could refer to using these corpora, or your own corpora on Sketch Engine.

In methodology, when using Sketch Engine to create own corpora

The texts, as obtained in their original format, will be stored securely (with two-factor authentication) in a dedicated folder on the researcher's OneDrive; for safety and integrity, only the researcher and their promotors will have secure access to this folder drive.

Subsequently, the texts will be converted into text files, and uploaded as a private corpus on the Sketch Engine platform (Kilgarriff et al. 2014; www.sketchengine.eu). The corpora were automatically processed and annotated using Sketch Engine's standard tools for English; see https://www.sketchengine.eu/corpora-and-languages/english-text-corpora/. Again, using Sketch Engine's secure functionalities, this corpus was shared only with the researcher and their promotors.

In acknowledgements

We acknowledge the generous financial support of the North-West University's (NWU) Faculty of Humanities for supporting the Corpus Cooperative at NWU (COCO@NWU). Through its aims to advance corpus-based research in the digital humanities, it directly supported and benefitted this research. Nonetheless, all searches, calculations, and interpretations are the researcher's only, and cannot be attributed to NWU.