Corpus Info

Available corpora

Corpora at NWU have mostly been developed in collaboration with various corpus suppliers, such as publishing houses, news websites, (literary) blog sites, libraries, etc. These corpora may be used for academic research purposes only. For commercial licencing options, please contact us directly.

These corpora are divided into two collections:

Available interfaces / platforms

COCO@NWU hosts the corpora on two platforms:

COCO@NWU Regular

Size
  • Total words: 409 889 362
  • Total tokens: 471 883 877
  • Frequency: "die": 30 228 084

A collection of government documents from the various websites and electronic publications of the South African government.

Size
  • Total words: 1 514 385
  • Total tokens: 1 736 282
  • Frequency: "die": 1 28 461
Reference

Department of Sports, Arts and Culture & CTexT. 2026. NCHLT- Afrikaanse Korpus 2.0 [NCHLT Afrikaans Corpus 2.0]. Potchefstroom: CTexT, North-West University. ISLRN: 544-932-849-161-3. Available at: https://hdl.handle.net/20.500.12185/293

A subset of the Leipzig Corpora Collection.

Subset contains Afrikaans texts up until 2023. © 2025 Abteilung Automatische Sprachverarbeitung, Universität Leipzig.

Size
  • Total words: 169 956 164
  • Total tokens: 193 682 495
  • Frequency: "die": 11 670 617
Reference

Goldhahn, D. Eckart, T. & Quasthoff, U. 2012. Building Large Monolingual Dictionaries at the Leipzig Corpora Collection: From 100 to 200 Languages. In: Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12).

A collection of original plays submitted for the ATKV Youth Theatre Competition.

Collection covers competition entries from 2016 to 2022.

Size
  • Total words: 4 007 970
  • Total tokens: 5 025 644
  • Frequency: "die": 155 470
Reference

ATKV & CTexT. 2026. NWU/ATKV-Tienertoneelkorpus 2.0 [NWU/ATKV Theatre for Teenagers Corpus 2.0]. Potchefstroom: CTexT, North-West University.

A collection of news articles and blogs as published on the media house Maroela Media's website.

Collection covers all material up to the end of August 2026.

Size
  • Total words: 63 121 535
  • Total tokens: 72 725 284
  • Frequency: "die": 4 524 239
Reference

Maroela Media & CTexT. 2026. NWU/Maroela Mediakorpus 3.1 [NWU/Maroela Media Corpus 3.1]. Potchefstroom: CTexT, North-West University.

A corpus of Afrikaans books (mostly fiction) published by Lapa Publishers.

Size
  • Total words: 19 981 313
  • Total tokens: 23 607 580
  • Frequency: "die": 908 282
Reference

LAPA Uitgewers & CTexT. 2026. NWU/LAPA-Korpus 2.0 [NWU/LAPA Corpus 2.0]. Potchefstroom: CTexT, North-West University.

A collection of articles published in the Taalgenoot magazine.

Collection covers all material from 2006 to 2025.

Size
  • Total words: 2 925 479
  • Total tokens: 3 422 587
  • Frequency: "die": 171 574
Reference

ATKV & CTexT. 2026. NWU/ATKV-Taalgenootkorpus 2.0 [NWU/ATKV Taalgenoot Corpus 2.0]. Potchefstroom: CTexT, North-West University.

A corpus of Afrikaans books (fiction and non-fiction) published by the publisher Protea Boekhuis.

Size
  • Total words: 14 093 645
  • Total tokens: 16 395 760
  • Frequency: "die": 959 753
Reference

Protea Boekhuis & CTexT. 2026. NWU/Protea Boekhuiskorpus 3.0 [NWU/Protea Boekhuis Corpus 3.0]. Potchefstroom: CTexT, North-West University.

A collection of news bulletins, as broadcast on Radio Sonder Grense and published on their website.
Collection covers all material from January 2005 to July 2026.

Size
  • Total words: 51 809 307
  • Total tokens: 57 115 881
  • Frequency: "die": 5 120 855
Reference

Radio Sonder Grense & CTexT. 2026. NWU/RSG-nuuskorpus 3.1 [NWU/RSG News Corpus 3.1]. Potchefstroom: CTexT, North-West University.

A stratified corpus used by the Afrikaans Language Commission, consisting of a variety of genres and domains, including academic journals, newspapers, literary works, informal writings, etc.

Size
  • Total words: 46 470 173
  • Total tokens: 53 643 030
  • Frequency: "die": 3 623 055
Reference

Taalkommissie van die Suid-Afrikaanse Akademie vir Wetenskap en Kuns. 2011. Taalkommissiekorpus 2.0 [Language Commission Corpus 2.0]. Potchefstroom: CTexT, North-West University.

A collection of documents from the various web pages available on the Afrikaans version of Wikipedia.
Collection covers all material up to July 2026.

Size
  • Total words: 32 159 649
  • Total tokens: 40 131 001
  • Frequency: "die": 2 732 615
Reference

Wikipedia. 2026. Wikipedia Afrikaans Corpus 2.1. Potchefstroom: CTexT, North-West University.

The data included in the corpora presented here issue from Pula/Imvula, a South African magazine focusing on the developing farmer and published by Grain SA (https://www.grainsa.co.za/farmer-development). The main aim of the magazine is to support developing farmers in becoming sustainable commercial farmers. The magazine is distributed on a monthly basis and currently only available in English. However, from 2007 until September 2024 it was published in five languages (English, isiXhosa, isiZulu, Sesotho, and Setswana), and previously (2007-2019) also included Afrikaans and Sesotho sa Leboa (discontinued due to lack of funding).

Size
  • Total words: 706 337
  • Total tokens: 798 067
  • Frequency: "die": 48 560
Reference

Pula Imvula & CTexT. 2025. NWU/Pula Imvula Afrikaans Corpus 1.0. Potchefstroom: CTexT, North-West University.

Klyntji.com is an independent, online journal published in all variants of Afrikaans, focusing on diverse and progressive art and culture from the global South. The corpus covers material from 2014 to 2026.

Size
  • Total words: 841 564
  • Total tokens: 969 581
  • Frequency: "die": 46 861
Reference

Klyntji & CTexT. 2026. NWU/Klyntji Corpus 1.2. Potchefstroom: CTexT, North-West University.

Vrye Weekblad was a progressive Afrikaans national weekly newspaper published from 1988 to 1994. It was relaunched as an online newspaper from 6 April 2019 to 28 March 2025. The corpus includes material from 2019 to 2025, but excludes user comments.

Size
  • Total words: 2 301 841
  • Total tokens: 2 630 685
  • Frequency: "die": 137 742
Reference

Vrye Weekblad & CTexT. 2025. NWU/Vrye Weekblad-korpus 1.1 [NWU/Vryeweekblad Corpus 1.1]. Potchefstroom: CTexT, North-West University.

Size
  • Total words: 1 028 794
  • Total tokens: 1 166 682
  • Frequency: "the": 65 232

The data included in the corpora presented here issue from Pula/Imvula, a South African magazine focusing on the developing farmer and published by Grain SA (https://www.grainsa.co.za/farmer-development). The main aim of the magazine is to support developing farmers in becoming sustainable commercial farmers. The magazine is distributed on a monthly basis and currently only available in English. However, from 2007 until September 2024 it was published in five languages (English, isiXhosa, isiZulu, Sesotho, and Setswana), and previously (2007-2019) also included Afrikaans and Sesotho sa Leboa (discontinued due to lack of funding).

Size
  • Total words: 1 028 794
  • Total tokens: 1 166 682
  • Frequency: "the": 65 232
Reference

Pula Imvula & CTexT. 2025. NWU/Pula Imvula English Corpus 1.0. Potchefstroom: CTexT, North-West University.

Size
  • Total words: 904 108
  • Total tokens: 1 014 905
  • Frequency: "go": 52 179

The data included in the corpora presented here issue from Pula/Imvula, a South African magazine focusing on the developing farmer and published by Grain SA (https://www.grainsa.co.za/farmer-development). The main aim of the magazine is to support developing farmers in becoming sustainable commercial farmers. The magazine is distributed on a monthly basis and currently only available in English. However, from 2007 until September 2024 it was published in five languages (English, isiXhosa, isiZulu, Sesotho, and Setswana), and previously (2007-2019) also included Afrikaans and Sesotho sa Leboa (discontinued due to lack of funding).

Size
  • Total words: 904 108
  • Total tokens: 1 014 905
  • Frequency: "go": 52 179
Reference

Pula Imvula & CTexT. 2025. NWU/Pula Imvula Sesotho sa Leboa Corpus 1.0. Potchefstroom: CTexT, North-West University.

Size
  • Total words: 1 235 476
  • Total tokens: 1 376 803
  • Frequency: "ho": 71 856

The data included in the corpora presented here issue from Pula/Imvula, a South African magazine focusing on the developing farmer and published by Grain SA (https://www.grainsa.co.za/farmer-development). The main aim of the magazine is to support developing farmers in becoming sustainable commercial farmers. The magazine is distributed on a monthly basis and currently only available in English. However, from 2007 until September 2024 it was published in five languages (English, isiXhosa, isiZulu, Sesotho, and Setswana), and previously (2007-2019) also included Afrikaans and Sesotho sa Leboa (discontinued due to lack of funding).

Size
  • Total words: 1 235 476
  • Total tokens: 1 376 803
  • Frequency: "ho": 71 856
Reference

Pula Imvula & CTexT. 2025. NWU/Pula Imvula Sesotho Corpus 1.0. Potchefstroom: CTexT, North-West University.

Size
  • Total words: 1 218 908
  • Total tokens: 1 335 498
  • Frequency: "go": 95 722

The data included in the corpora presented here issue from Pula/Imvula, a South African magazine focusing on the developing farmer and published by Grain SA (https://www.grainsa.co.za/farmer-development). The main aim of the magazine is to support developing farmers in becoming sustainable commercial farmers. The magazine is distributed on a monthly basis and currently only available in English. However, from 2007 until September 2024 it was published in five languages (English, isiXhosa, isiZulu, Sesotho, and Setswana), and previously (2007-2019) also included Afrikaans and Sesotho sa Leboa (discontinued due to lack of funding).

Size
  • Total words: 1 218 908
  • Total tokens: 1 335 498
  • Frequency: "go": 95 722
Reference

Pula Imvula & CTexT. 2025. NWU/Pula Imvula Setswana Corpus 1.0. Potchefstroom: CTexT, North-West University.

Size
  • Total words: 712 468
  • Total tokens: 825 364
  • Frequency: "ukuba": 10 398

The data included in the corpora presented here issue from Pula/Imvula, a South African magazine focusing on the developing farmer and published by Grain SA (https://www.grainsa.co.za/farmer-development). The main aim of the magazine is to support developing farmers in becoming sustainable commercial farmers. The magazine is distributed on a monthly basis and currently only available in English. However, from 2007 until September 2024 it was published in five languages (English, isiXhosa, isiZulu, Sesotho, and Setswana), and previously (2007-2019) also included Afrikaans and Sesotho sa Leboa (discontinued due to lack of funding).

Size
  • Total words: 712 468
  • Total tokens: 825 364
  • Frequency: "ukuba": 10 398
Reference

Pula Imvula & CTexT. 2025. NWU/Pula Imvula isiXhosa Corpus 1.0. Potchefstroom: CTexT, North-West University.

Size
  • Total words: 1 352 133
  • Total tokens: 1 570 257
  • Frequency: "futhi": 12 212

The data included in the corpora presented here issue from Pula/Imvula, a South African magazine focusing on the developing farmer and published by Grain SA (https://www.grainsa.co.za/farmer-development). The main aim of the magazine is to support developing farmers in becoming sustainable commercial farmers. The magazine is distributed on a monthly basis and currently only available in English. However, from 2007 until September 2024 it was published in five languages (English, isiXhosa, isiZulu, Sesotho, and Setswana), and previously (2007-2019) also included Afrikaans and Sesotho sa Leboa (discontinued due to lack of funding).

Size
  • Total words: 661 442
  • Total tokens: 774 127
  • Frequency: "futhi": 8 489
Reference

Pula Imvula & CTexT. 2025. NWU/Pula Imvula isiZulu Corpus 1.0. Potchefstroom: CTexT, North-West University.

A collection of news stories published by KZN Namuhla, a Community Newspaper that serves the Durban Metropolitan semi-urban and rural areas, Pietermaritzburg as well as the South and North of KZN. A few opinion pieces are also included.
Collection covers all material from October 2022 to July 2026.

Size
  • Total words: 690 691
  • Total tokens: 796 130
  • Frequency: "futhi": 3 723
Reference

KZN Namuhla & CTexT. 2026. NWU/KZN Namuhla News Corpus 1.0. Potchefstroom: CTexT, North-West University.

COCO@NWU Special

Size
  • Total words: 72 690 022
  • Total tokens: 83 549 373
  • Frequency: "die": 2 976 186

Collection of unedited user comments on web pages, arranged in subcorpora per web page.

Current collection: Comments from two different websites (anonymous).

Size
  • Total words: 51 498 974
  • Total tokens: 58 783 978
  • Frequency: "die": 2 819 831
Reference

CTexT. 2026 . NWU Kommentaarkorpus 3.1 [NWU Comment Corpus 3.1]. Potchefstroom: CTexT, North-West University.

Collection of historical Afrikaans documents, arranged by subcorpus.

Current collection: VOC documents from 1662 to 1791, as digitized by the Tracing History Trust.

Size
  • Total words: 19 343 422
  • Total tokens: 22 647 357
  • Frequency: "die": 67 177
Reference

Tracing History Trust & CTexT. 2021 . NWU THT-korpus 1.4 [NWU THT corpus 1.4]. Potchefstroom: CTexT, North-West University.

Collection of informal blogs as published on watkykjy.co.za up to and including the end of February 2025.

Size
  • Total words: 1 847 626
  • Total tokens: 2 118 038
  • Frequency: "die": 89 178
Reference

WatKykJy & CTexT. 2026. NWU/WatKykJy-korpus 3.0 [NWU/WatKykJy Corpus 3.0]. Potchefstroom: CTexT, North-West University.

Automatic corpus processing

For both search platforms, all non-English corpora have been automatically processed, annotated, and (re)compiled using in-house tools developed by CTexT. The ctextcore Python package was used for sentence separation, tokenisation and POS annotation.

NOTE: The Leipzig corpus was created by combining all available Afrikaans Leipzig collections by year and type. Deduplication was performed at the paragraph level for each combined collection.

Processing and annotation of Afrikaans texts includes the following levels:

Afrikaans Sentence separation and tokenisation

Note that Sentence separation and tokenisation counts might differ slightly between Afrikaans corpora that users have uploaded themselves on Sketch Engine, and those that have been uploaded by NWU on the Sketch Engine and COCO platforms. The main reason for this is that Sketch Engine and our in-house tools use slightly different definitions of what sentences and words are.

POS tagsets

The POS tagsets for all availabe languages can be accessed at: Annotation Tag Sets

Citations and references

Initiative/website

COCO@NWU. 2025. Corpus Cooperative at North-West University. Available at: http://coco.nwu.ac.za.

Corpus collections

Choose between one or more of the following:

Annotation / processing tools

See the citation information under each of the tools mentioned above.

Sample texts

You may use / adapt the texts below in your research proposal, ethics application, article, thesis, etc. They merely serve as examples how you could refer to using these corpora, or your own corpora on Sketch Engine.

In methodology, when using Sketch Engine to create own corpora

The texts, as obtained in their original format, will be stored securely (with two-factor authentication) in a dedicated folder on the researcher's OneDrive; for safety and integrity, only the researcher and their promotors will have secure access to this folder drive.

Subsequently, the texts will be converted into text files, and uploaded as a private corpus on the Sketch Engine platform (Kilgarriff et al. 2014; www.sketchengine.eu). The corpora were automatically processed and annotated using Sketch Engine's standard tools for English; see https://www.sketchengine.eu/corpora-and-languages/english-text-corpora/. Again, using Sketch Engine's secure functionalities, this corpus was shared only with the researcher and their promotors.

In acknowledgements

We acknowledge the generous financial support of the North-West University's (NWU) Faculty of Humanities for supporting the Corpus Cooperative at NWU (COCO@NWU). Through its aims to advance corpus-based research in the digital humanities, it directly supported and benefitted this research. Nonetheless, all searches, calculations, and interpretations are the researcher's only, and cannot be attributed to NWU.