Morphology as linguistic discipline
This topic provides an overview of morphology as a linguistic discipline. Hereafter a brief description of the morphological processes involved with word formation is presented. This background of general morphology forms the basis for a discussion on the typological characteristics of Setswana as they relate to word categories and the tokenisation of Setswana words.
Morphology as linguistic discipline
Morphology is the linguistic study of morphemes and their variants/allomorphs and is specifically concerned with the way in which morphemes function in the formation of words (Louwrens, 1994: 115). In other words, morphology is the study of the internal structure of words (Kosch, 2006: 1) and it has close links with the other linguistic disciplines. The term morphophonology indicates its close link with phonology which is the study of the function attached to sound. The term morphosyntax indicates its close link to syntax which is the study of how words combine to form phrases and sentences. It is clear from these terms that the borders between morphology, phonology and syntax are blurred. All three these disciplines are also highly involved with the meaning portrayed by the linguistic unit referred to, semantics.
The grammatical units that set up a language are morphemes and words which are distinguished in morphology, and phrases, clauses and sentences which are distinguished in syntax. Grammar is the term that is conventionally used for the rules that govern the formation and usage of these terms.
The purpose with morphology is to study and describe the structure of words by examining the meaningful sound sequences that compose these words. A morpheme is the smallest sound sequence which still has a meaning.
The word as unit in morphological description
The word is accepted as the unit that is focused on in morphological analysis. However, in universal grammar or parallel grammar this may not be the ideal basis for description as the linguistic units employed to express the same meaning may differ across languages. Words and morphemes in English for instance, alter when the same meaning is translated into Setswana. In the English sentence I will help you five words appear, each with its own meaning within the sentence. The equivalent meaning in Setswana would be:
| Ke tla go thusa |
| Ke -tla- go- thus -a |
| AgrSubjP1sg -FutPre- AgrObjP2sg- help -VEnd |
This would constitute one linguistic word i.e., a verb, consisting of four units in writing: Ke (‘subject agreement marker morpheme’), tla (‘future tense morpheme’), go (‘object agreement morpheme’), thus- ‘help’ and -a (‘verb ending/verb final suffix’).
The term word is thus used to refer to different concepts, on the one hand it is used to refer to the written word, a unit written separated from other units by white space. This is purely a matter of convention in the formalisation of the written representation. Such a word would be referred to as an orthographic word – a sound or sequence of sounds separated from other sounds or sequences of sounds by means of spaces in writing.
On the other hand, it may be used to refer to words that belong to the same word category (lexical category, word class, part of speech). This interpretation would refer to linguistic words – a unit which has its own independent meaning. Structurally it contains at least a root. In addition to the root it may contain another root (or roots) and/or affix(es) (Kosch, 2006: 3).
Setswana follows a disjunctive writing convention in which a single linguistic word may be represented by several orthographically separated units. This pertains mainly to the prefixes of the verb.
In this section, the term word is to be understood as the linguistic word. Linguistic words are classified into groups referred to as word categories (lexical categories or parts of speech). Setswana words are classified into eight word categories viz. nouns, pronouns, verbs, particles, conjunctions, adverbs, interjections, and ideophones. Some of these word categories can in turn, be divided into subcategories (Van Wyk, 1966; Krüger, 2006). It is customary to distinguish between open and closed word categories in the Bantu languages. Nouns and verbs are open (morphologically productive) categories while pronouns, particles, conjunctions, adverbs, interjections and ideophones constitute the closed (morphologically unproductive) categories (Pretorius et al., 2008: 2; Pretorius, 2014: 49).
Words differ regarding morphological structure, as some words may consist of a single morpheme and others may have a complex structure. Monomorphemic words are written as one orthographic unit and cannot be divided into smaller units, for example the Setswana adverb ruri ‘really, truly’, tota ‘really, truly’, gape ‘again, also’ or Setswana conjunctions such as fa ‘if, when’, ka ‘as’:
| Fa ke di bona (dinotlolo) |
| Fa ke di bon-a (dinotlolo) |
| Conj AgrSubjP1sg AgrObj10 see-VEnd (keys) |
| If I find(see) them. |
| ke a di batla |
| ke-a-di-batl-a |
| I am looking for them. |
The affixes ke-, a-, di- and -a are attached to the root batl-. Combined these morphemes constitute the word ke a di batla, a verb. A single linguistic word is thus represented by several orthographic words (cf. Creissels et al., 1997).
The morpheme as the smallest/minimal unit
The morpheme (from the Greek word morph - form) is defined as the smallest minimum meaningful part of a word-form, in other words, as the smallest indivisible unit of meaning. A morpheme is the smallest meaning bearing unit in a language, the minimal distinctive unit in the grammar of a language and the main focus or central concern of morphology. Katamba (1993: 25) gives this definition: “The morpheme is the smallest difference in the shape of a word that correlates with the smallest difference in word or sentence meaning or in grammatical structure”.
To recognise the morphemes of Setswana it is necessary to identify the smallest units of sounds which constantly present a specific meaning and function in collections of words. Consider the words in the following table:
| A word | B meaning | C word | D meaning |
|---|---|---|---|
| setlhare | ‘tree’ | ditlhare | ‘trees’ |
| sekgwa | ‘forest’ | dikgwa | ‘forests’ |
| setlhako | ‘shoe’ | ditlhako | ‘shoes’ |
| seboko | ‘worm’ | diboko | ‘worms’ |
| seatla | ‘hand’ | diatla | ‘hands’ |
All the words in column A start with se- which indicates that the difference in meaning between the words in column A has to be in -tlhare ‘tree’, -kgwa ‘forest’, -tlhako ‘shoe’, -boko ‘worm’ and -atla ‘hand’. The meaning/function of se- is recognised when the words in column A are compared to their counterparts in column D. It becomes evident that se- indicates the singular form of the words in column A while di- indicates the plural form of the words in column C. Both se- and di- thus indicate number and can be taken to be morphemes. The other parts of the words -tlhare, -kgwa, -tlhako, -boko and -atla appear in both columns A and C with the same meaning and can therefore also be taken to be morphemes. Dividing the words in any other way such as setlh-, ar- and -e will not render a morphologically viable result. Morphemes have to consist of the smallest possible number of sounds to which a specific meaning can be ascribed. This meaning may be independent or a part of a meaning or even represent a certain grammatical category as shown for se- and di- above. A morpheme is therefore the smallest meaning bearing unit in a language. Morphemes are classified into free morphemes (i.e., words) and bound morphemes.
Note that in some Setswana dialects such as Sengwaketse, leaving out the prefixes is acceptable.
The classification of morphemes
Linguists in general and those working in the Bantu languages as such equate the notion word with a root or simplex, thereby neutralising the distinction between morpheme and word on the one hand and ignoring complexes, which can also be parts of words, on the other hand.
It is important to acknowledge the existence of intermediate structures in the morphological composition of Setswana words. Motse ‘town’ is an intermediate structure in motseng ‘in, to, from the town’ as monna ‘man’ is an intermediate structure in monnanyana ‘small / little man’. The parts motse and monna are meaningful parts that are qualitatively totally different from -ng and -nyana. To accommodate these items in the hierarchical layers of the word form, terms such as root, stem, base, lexical morpheme, free and bound morpheme, etc. have been created.
Free and bound morphemes
Morphemes can be free or bound. Spencer (1991: 5) describes the relationship between word and morpheme as:
“The fact that one and the same entity can be both a morpheme and a word (or, equivalently, that some words consist of just one morpheme, i.e., are monomorphemic) shouldn't worry us. However, it is useful to distinguish those morphemes that are also words in their own right from those that only appear as a proper subpart of a word. The former are called free morphemes and the latter bound morphemes.”
Thus, morphemes that can operate independently as words are referred to as free morphemes while morphemes that can only appear as a part of a word are referred to as bound morphemes.
In Setswana monomorphemic words such as some adverbs like tota ‘really / truly’ and ruri ‘really, truly’ and gape ‘again, also’ as well as particles like fa ‘if, when’ and ka ‘as’ can be classified as free morphemes. In a polymorphemic word such as setlharenyana or commonly setlhatshana ‘a small tree’ the morphemes se- (class prefix which marks class number), -tlhare- ‘tree’ and -nyana (diminutive suffix) are all bound morphemes.
Roots and stems
Earlier views in grammars of South African African languages regarding morphological components and morphological analysis stemmed mainly from the work of Doke who was a Zulu scholar. Doke, as quoted by Ziervogel et al. (1985: 287) states that:
“The distinction between roots and stems is more or less arbitrary, and one employed for convenience ... In fact, a stem is, generally speaking, that part of the word which is shorn of its prefixal elements. Take, for instance, the stems -thanda, -thandisa, -thandana of the verb, and -thando of the noun.”
Examples for Setswana would be: -batla; -batlisa; -batlana of verb; and patlo of noun.
Lombard et al., (1985: 24) in following Doke defines the stem as:
“The stem of a word consists of the root plus all the suffixal morphemes in the word. The stem is distinguished as a structure level on the grounds of the fact that it functions as a unit in the tonology (or tone morphology). When a root has no suffixal morphemes we therefore do not speak of a stem either, but of a root only.”
Consequently, in a word such as motsamaisi ‘colleague’ the stem would be -tsamaisi* and in mosadinyana ‘small woman’ the stem would be -sadinyana*. The objection to this view is obvious, as both of these segments are unsuited and unusable for linguistic (morphological and syntactic) description. They also do not have an independent meaning.
Posthumus (1994: 14) questions the Dokean approach:
“Because suffixing is a far more productive morphological process in the languages of the world, the term ‘stem’ has often been defined in terms of suffixes. Why should the root plus any suffix(es) be called a stem? What morphological significance is vested in the suffixes of the African languages not contained in the prefixes? There is no justification for such a definition in the African languages.”
He then illustrates it by way of examples in Zulu:
“Why should in ba-dlile ‘they have eaten’ the form -dlile be viewed as the stem and not ba-dl(a) which is anyway the basic underlying form? Meanwhile it should also be noted that -dlile is non-existent.”
Krüger (2006: 45) agrees with Posthumus (1994) and encourages a hierarchical analysis of pre- and suffixes when he states that:
“a hierarchical analysis implies that the underlying components must themselves be meaningful and not meaningless or non-existent. For example to analyze a word such as ungrateful into ungrate and ful or moagi ‘builder’ into mo- and -agi would be wrong because ungrate and -agi are meaningless and non-existent and they can in no way whatsoever function as underlying forms.”
He goes on to discuss and explain the hierarchical method of removing grammatical morphemes from words in this fashion until only the root is arrived at.
Pretorius (2000) provides a detailed outline of opinions on the terms root and stem in the Sotho languages.
Root
Roots are morphemes that cannot be omitted in words as they carry the basic lexical meaning of the word.
In Setswana words such as setlhare ‘tree’ or settlharenyana/setlhatshana ‘small tree’ or setlhareng ‘in (to/at) the small tree’ the root is -tlhare, and it appears in all the derivations of the word. However, in noun class 9 nouns such as dikatse ‘cats’, katsenyana/katsana ‘small cat’ or katseng ‘(at/on) the cat’ the segment katse is not a root but a stem because it can function as an independent word. The root -tlhare above cannot function as a word.
Stem
A stem is a lexical morpheme and it is that part of a word-form which:
- may include grammatical morphemes,
- can operate independently,
- constitutes the lexical meaning of a word,
- belongs to an open class and
- has a word-correlate outside the structure in which it occurs.
The key difference between a root and a stem is that a stem is a word - it has a word-correlate; a root cannot function as a word.
Stems can be classified into the following categories on the grounds of their morphological composition:
a) Simple stem (Simplex stem/Monomorphemic stem)
| Lexical morpheme | Translation | Word | Translation | |
|---|---|---|---|---|
| ba | ‘these’ | in | bao | ‘those’ |
| ntate | ‘father/sir’ | in | bôntate | ‘father and company’ |
| koko | ‘chicken’ | in | kokwana | ‘small chicken’ |
b) Complex stem (Complex stem/Polymorphemic stem)
| Lexical morphemes | Translation | Word | Translation | |
|---|---|---|---|---|
| monwana | ‘finger/toe’ | in | monwananyana | ‘little/small finger’ |
| monwananyana | ‘little/small finger’ | in | monwananyanêng | ‘in, on, from the little/small finger’ |
c) Reduplication stem (Duplicated stem)
| Stem | Word | Translation | |
|---|---|---|---|
| motsemotse | in | motsemotsana | a tiny little village |
| mosadisadi | in | mosadisadinyana | a real small little town |
| Example | Translation |
|---|---|
| go-kwalakwala | ‘to write insignificantly, now and then’ |
| go-emaema | ‘to stand around aimlessly’ |
d) Compound stem
| Compound stem | Seperate stems |
|---|---|
| leebarope (‘rock pigeon’) | leeba, marope (‘pigeon’, ‘ruins’) |
| pelonolo (‘kindness’) | pelo, bonolo (‘heart’, ‘gentleness’) |
| monnamogolo (‘old man’) | monna yô mogolo (‘man that is old’) |
Base / operand / underlying component
The term base has a specific application in the description of morphological operations. The linear arrangement of morphemes (morpheme syntax) in the formation of words is determined by a number of morphological operations. When these operations are functional it is essential to determine the characteristics of the base from which is operated also referred to as the operand (Matthews, 1989: 124). The base or operand acts as the underlying or immediate component when more complex meaningful forms, referred to as the derivand, are established by means of these operations. (Posthumus, 1994: 31; Krüger, 1994: 20).
In the hierarchical approach followed by the above scholars it is important for both the operand and the derivand to be meaningful. Consider the following examples:
| sekolonyanêng |
| se- -kolo- -nyana -ing |
| NPre school- DimSuf LocSuf |
| in, from, to the little school |
In the morphological analysis of the polymorphemic stem (kwa) sekolonyanêng ‘(at) the little school’ it is possible to identify several parts which may act as the immediate operands (bases) from which derivands are formed:
| Base | Derivands |
|---|---|
| -kolonyanêng | to which se- is attached |
| -kolong- | to which se- and -nyana are attached |
| sekolo- | to which -nyanêng is attached |
| sekolonyana- | to which -(i)ng is attached |
The options -kolonyanêng, and -kolong are meaningless and non-existent entities in the same way that unfriend* is meaningless in the analysis of the English adverb unfriendly. Sekolo- ‘school’ and sekolonyana- ‘little school’ are meaningful parts of the more complex words: sekolo- in sekolonyana-, and sekolonyana- in sekolonyaneng. Sekolo is thus the operand for sekolonyana while sekolonyana is the operand for sekolonyaneng.
The example and explanation above should serve as motivation that a word-based approach to morphology should be opted for when performing morphological operations that are part of word formation and analysis. The importance of meaningful operands as starting points instead of meaningless non-existent “parts”
of words is logical. For analysis it implies that every underlying component should be a word-correlate of an existing word outside the structure. For word-forming operations it implies that only existing word-correlates should be used as bases (operands) and not meaningless segments (Katamba, 1993).
Grammatical morphemes / affixes
Grammatical morphemes are affixes that indicate various grammatical categories such as class, number, aspect, tense, consecutive, passive, applicative, etc. These ‘meanings’ can be defined as grammatical or semantic values. Based on their influence on the word, grammatical morphemes can be divided into 2 categories, i.e., inflectional morphemes and derivational morphemes.
Inflectional morphemes
Inflectional morphemes are morphemes which do not modify, alter or change the lexical meaning of a word. They establish grammatical categories (values) in a word. The negative morpheme ga- and ending -e in Ga re go utlwe ‘we do not hear you’. The singular morpheme se- and the plural morpheme di- in setlhare ‘tree’ and ditlhare ‘trees’ are examples of inflectional morphemes.
Derivational morphemes
Derivational morphemes modify the lexical meaning of a word. In the verb go ja ‘to eat’ the causative morpheme -is- would introduce the meaning ‘to feed/to let eat’ in go jesa, while the passive morpheme -iw- would introduce the meaning ‘to be eaten’ in go jewa. In a noun such as sekolonyaneng ‘in the little school’ the diminutive morpheme -nyana and the locative morpheme -ng are derivational morphemes.
Classification of morphemes according to distribution (position)
Morphemes can also be classified according to their relevant positions in the word. Grammatical morphemes are positioned on the periphery, to the left and/or to the right of roots. Prefixes are affixes that precede the root of the stem that they appear in. Suffixes are affixes which follow the root of the stem that they appear in. Roots are the stable points around which the grammatical morphemes are distributed and could be referred to as central morphemes.
Morphological processes in word formation
The two central phenomena in morphology are word formation (also referred to as morpheme sequencing or morphotactics) and phonological and orthographical alternation (also referred to as morphophonological alternation) – the sound and spelling changes that occur due to the environment in which a morpheme occurs. In Setswana, as an agglutinative language, both these phenomena play an important role (Berg, 2018: 3). Affixes are sequenced as structural elements in a word to execute a process of adapting or extending the meaning of a word (Kosch, 2006: 133-139).
Derivation and inflection are conventionally known as the two most important types of morphological process. Morphemes which change a word on the semantic level are called derivational morphemes/affixes, the morphological process is known as derivation. Inflectional affixes specify the grammatical functions of words in phrases without altering their meaning. The meaning of a noun can be extended by a diminutive, feminine, augmentative and locative suffix (Krüger, 2006: 73-96). Inflection in verb morphology is expressed by prefixes that indicate class gender, person and number, mood, tense, aspect, and polarity (Cole, 1955: 242-267; Krüger, 2006: 198-243).
The difference between derivation and inflection lies mainly in the fact that in derivation the lexical content of a word is affected, while in inflection it is not.
When more complex words are formed from less complex words certain techniques are formed. These techniques are known as word-formation processes / operations. For Setswana the following techniques can be mentioned, i.e., affixation, substitution, reduction, reduplication and compounding.
Affixation
Affixation is a very productive word forming technique in Setswana. It is the process whereby affixes (grammatical morphemes) are prefixed and/or suffixed to the operand (underlying component) to form more complex words. In an example such as go tsidifatsa ‘to make cold / to refrigerate’ the infinitive prefix go- is prefixed to the tsididi concomitantly with the suffixing of the denominative suffix -fal- and ending -a to tsididi to form go tsidifala ‘to become cold’. This would then form the operand or underlying component to which the causative suffix -y- is suffixed to render go tsidifatsa ‘to refrigerate/to make cold’.
In a noun such as koloinyana ‘small vehicle’ the diminutive suffix -nyana is suffixed to koloi, the operand. In koloinyaneng ‘in/at the small vehicle’ the locative suffix -(i)ng is suffixed to koloinyana.
Substitution
Substitution is also a productive technique for word formation in Setswana. It is a process whereby one morpheme is replaced by another in order to change a grammatical category such as the number (singular or plural) of a noun or the polarity (positive or negative) of a verb (Louwrens, 1994: 190).
In mosadi ‘woman’ > basadi ‘women’ (‘plural’) the singular prefix mo- is replaced by the plural prefix ba-.
The positive polarity of the verb ba a opela dipina ‘they are singing songs’ changes to negative polarity by prefixing the negative morpheme ga- to ba a opela and concomitantly substituting the final vowel -a with -e and ga ba opele ‘they are not singing’ is then formed.
Reduction
Reduction is a word formation technique whereby a morpheme is removed or taken away from a (more complex) word in the formation of another word. Reduction is not a common proses in Setswana. A case in point may be the formation of imperatives such as Dula! ‘Sit (down)!’ where the infinitive prefix go in go dula ‘to sit (down)’ is removed.
Reduplication
Reduplication is a term that is used to refer to the process whereby a word or a part of a word is duplicated/repeated. This is done to achieve a particular semantic effect such as emphasis or to convey the idea that an action is carried out intermittently (Louwrens, 1994: 163). Full or partial duplication may appear.
The reduplication of nouns indicates intensification:
| godimodimo |
| at the very top |
| tlasetlase |
| at the very bottom |
Reduplication of verbs indicates that the action is carried out intermittently:
| go tabogataboga |
| to run every now and then (insignificantly) |
| go balabala |
| to speak repeatedly or intermittently |
Reduplication as process of word formation is not as productive as processes mentioned above.
Compounding
The principal devices for the formation of words are derivation and compounding. Compounding is closely related to reduplication. Compounding is a process whereby two (or more) autonomous words, mostly members of a word group, are combined to form a compound word:
| monnamogolo |
| an old man / big man |
| mo- nna mo- golo |
| NPre1-man NPre1-big |
| khudutlou |
| a very large tortoise |
| (ne-)-khudu (ne-)-tlou |
| NPre9-tortoise NPre9-elephant |
Compounding is a very productive technique or process in the coining of new Setswana words particularly with regard to the coining of new terms.
Setswana word categories
Setswana words are classified into eight word categories (parts of speech, lexical categories, word classes) viz. nouns, pronouns, verbs, particles, conjunctions, adverbs, interjections and ideophones (Berg, 2018: 47). Some of these word categories can in turn be divided into subcategories (Van Wyk, 1966; Krüger, 2006). Interrogatives are not identified as a separate word category in Setswana. Krüger (2013: 349-371) describes the word status and function of various Setswana interrogatives.
Setswana word categories are often grouped into open and closed word categories based on morphological productivity. Nouns and verbs are open (morphologically productive) categories while pronouns, particles, conjunctions, adverbs, interjections and ideophones make up the closed (morphologically unproductive) categories (Pretorius et al., 2008: 2; Pretorius, 2014: 49).
Krüger (2013: 349-371) indicates that the interrogatives eng? ‘what?’, kae? ‘where?’, leng? ‘when?’ and jang? ‘how?’ are generally classified as adverbs in Setswana because of their usage in typical adverbial position. Khoali (1994) gives an exposition of the use of the different Setswana interrogatives.
Tokenisation and tagging of Setswana words
Tokenisation is the segmentation of running text into tokens such as words, numbers, punctuation marks, parentheses and similar entities (Berg, 2018: 160). Tokenisation is the process of breaking the text down into subunits (Grefenstette, 1999: 117), those strings of characters in texts that receive individual tags (Van Halteren & Voutilainen, 1999: 110). A tag is a grammatical class, embodied in some annotation scheme, which is associated with a word in a text. For morphology, tokens that are linguistic words are paramount. The tokenisation and tagging of linguistic components are important pre-processing steps in Human Language Technology (generally referred to as HLT), which is the computational treatment of language via electronic resources.
The difference between an orthographic and a linguistic word in Setswana is discussed in section 1. This is an important difference for Setswana tokenisation since the verb usually consists of multiple orthographic words but only one linguistic word. Refer also to Creissels et al. (1997)
Tokens to be considered as linguistic words are discussed in Pretorius (2014) while a detailed discussion of Setswana word categories and their sub-categories as well as a general-purpose part of speech tagset is available in Van Rooy & Pretorius (2003).
Typological characteristics of Setswana
The typology of languages can be treated with regard to various features of the language:
- Phonetic typology - the ranges of the sounds of the language.
- Phonological typology - ways in which sounds and sound features are organised into phonological systems and syllable structures.
- Grammatical typology - grammatical systems to mark syntactic relationships and sentence structure, for example, by word order and word class membership.
- Morphological typology - patterns that are typical in word-structure.
In this brief introduction the focus is on selected criteria in morphological description of the Bantu languages of which Setswana is one. Languages in their entirety cannot be neatly pigeonholed into a given class, the matter being a question of tendency. It is the prevailing characteristics that determine the basic type of language (Kosch, 2006: 132).
Setswana belongs to the Bantu language family and is classified in the South-Eastern Zone of Bantu languages, that is Zone S (Guthrie, 1967 – 1971). The South-Eastern Bantu languages are grouped together in language groups based on their similar grammatical structure and vocabulary (Poulos & Louwrens, 1994: 2; Krüger, 2006: 3).
Bantu languages are structurally closely related in terms of typology, as they share certain general characteristics such as a noun class system, a system of grammatical (concordial) agreement, and an agglutinative morphology (Louwrens, 1994: 18). Bantu languages differ regarding orthography. The difference is conventional rather than linguistically motivated. The Nguni languages have a conjunctive orthography in which affixes are conjoined with the root while the Sotho languages employ a disjunctive orthography in which the prefixes of the verb are generally written disjunctively.
The Bantu languages are characterised by a grammatical gender, so-called class gender, where nouns are grouped together in classes in a grammatically significant way (Kosch, 2006: 89-90). The nouns are grouped in classes by means of their class prefixes which are correspondingly referred to as gender number prefixes (Kosch, 2006: 90). Even though the numbering of noun cases is arbitrary it has been adopted for all the Bantu languages. Nouns without an evident prefix are classed according to their agreement. Moreover, Setswana noun classes have semantic significance (Cole, 1955: 68-105; Krüger, 2006: 57-98). Each Setswana noun belongs to one of 20 noun classes and are numbered systematically. Classes 1 to 14 consist of singular-plural pairs, noun classes 1 and 2, 3 and 4, 5 and 6, 7 and 8, and 9 and 10 are pairs where the odd numbers indicate the singular and the even numbers the plural. Nouns in class 11 are singular and their plural forms conform to class 10. The nouns in class 14 are singular but their plural forms conform to class 6. Classes 1 and 2 each have a sub class, i.e., classes 1a and 2a. Nouns in class 1a are singular and their plural counterparts appear in class 2a. Classes 15 to 20 do not denote singular or plural. Class 15 contains infinitive nouns. Classes 16 to 20 contain locative classes (Krüger, 2006: 92-98; Berg, 2018: 47).
Grammatical (concordial) agreement in the Bantu languages is based on the noun class system (Lombard et al., 1985: 54) which also indicates person and number. The noun class agreement system does not mark gender in Bantu languages.
Agglutinating characteristics of Setswana
The term 'agglutinating' derives from the Latin gluten ‘glue’. In the context of language, agglutination refers to a process whereby various affixes are 'glued' on or simply 'stuck' onto other morphemes in sequence (Kosch, 2006: 134). An agglutinating language is one in which there are a number of obligatorily bound morphs each of which realises in a single morpheme. That is, there is a one-to-one correspondence between morph and morpheme in such languages. This implies that ideally there are no allomorphs in such languages (Kosch, 2006: 133).
Typical features of agglutinating languages are:
a) Words typically have several morphemes; that is, they are polymorphemic and are often morphologically complex. Compare the following examples:
| moagisani |
| mo- ag- is- an- i |
| NPre1-build-CausSuf-RecSuf-DevSuf |
The Setswana noun: moagisani ‘neigbour’ is a deverbative noun consisting of:
- mo- the class prefix of noun class 1
- -ag- root of the verb (‘build’)
- -is- causative suffix
- -an- reciprocal suffix
- -i deverbative ending indicating a person, or an agentive marker
| o ba kwalisile |
| o- ba- kwal- is- il- e |
| AgrSubj1-AgrObj2-write-CausSuf-PerfSuf-VEnd |
The Setswana verb: Morutabana o ba kwadisitse teko. ‘The teacher had them write a test.’ The verb o ba kwadisitse consists of:
- o- subject agreement morpheme of noun class 1
- -ba- object agreement morpheme of noun class 2
- -kwal- root of the verb (‘write’)
- -is- causative suffix
- -il- perfect suffix
- -e verb ending / verb final suffix
The morphemes in the words above cannot function in isolation and must be part of a word.
b) Morphemes ideally undergo no changes and hence the boundaries between their morphs are clear-cut. Compare the following:
| Batho ba a thusana. |
| People help each other. |
| ba- tho ba- a- thus-an- a |
| NPre2-human AgrSubj2-PresPre-help-RecSuf-VEnd |
c) Individual grammatical categories or a semantic content may be fairly easily assigned to morphs, which follow one another in a specific order. This is also known as the principle of “one form – one meaning”. Affixation is used expansively in order to express a variety of meanings and grammatical relations. Each successive morph corresponds to a single morpheme. The morph, being a realisation of a morpheme, will express the same meaningful content (be it lexical or grammatical) as the morpheme which it represents. Each morpheme (meaning) can in turn be expected to be realised by just one morph. This can be described as a one-to-one relationship between the morph and the morpheme (Kosch, 2006: 135). Compare the following:
| Mokatisi o ba tabogisitse. |
| The coach had them run far. |
| o- ba- tabog-is- il- e |
| AgrSubj1-Agrobj2- run- CausSuf-PerSuf-VEnd |
The verb o ba tabogisitse consists of:
- o- subject agreement morpheme
- -ba- object agreement morpheme
- -tabog- root of the verb ‘write’
- -is- causative suffix
- -il- perfect suffix
- -e verb ending / verb final suffix
Isolating characteristics of Setswana
There are two pertinent features of isolating languages which can be observed in the Bantu languages and in Setswana. They are:
a) The occurrence of monomorphemic words
Even though polymorphemic words are the most common words in Setswana monomorphemic words occur in the following word categories:
| Word category | Examples |
|---|---|
| Class 9 polysyllabic nouns | kgomo ‘cow’, pene ‘pen’, koko ‘chicken’ |
| Nouns of class 1a | ntate ‘father’, malome ‘uncle’ |
| Conjunctions | mme ‘but’, gonne ‘because’ |
| Adverbs | fela ‘only/merely/just’, tota ‘truly/properly’ |
| Interjections | êê! ‘yes’, ao! ‘expressing surprise/reproach’ |
| Ideophones | tu/tuu ‘silence’, thô ‘liquid falling in large drops’ |
b) The use of tone for grammatical and lexical distinctions
Setswana also relies on tone for grammatical and lexical distinctions. (cf. Cole, 1975: 53-55; Creissels et al., 1997)
| mosimanyana |
| a little boy |
| mosimanyana |
| a little hole |
| se rêkê! |
| do not buy |
| se rêkê! |
| buy it |
Fusional characteristics of Setswana
In historical linguistics and phonetics, fusion or coalescence, is regarded as a sound change where two or more segments with distinctive features merge into a single segment. A fusional language is thus a language where there is no clear-cut boundary between morphs as different categories are fused together and expressed within the same unsegmentable morph. Setswana shows the following characteristics of fusional languages:
a) Variance of morphemes and fusion of morphs
According to Kosch (2006: 140):“A morpheme does not always appear with a constant phonological form, but often displays allomorphy. Furthermore, the different morphs constituting a word may not always be clearly segmentable. The boundaries between morphs become obscured when a root or stem of a word is modified (with or without some affixation) to reflect or accommodate a grammatical idea. In more drastic modifications, such as instances of suppletion, a morpheme assumes a totally different form to express added grammatical information”
Compare the following Setswana verbs:
| Bana ba robetse |
| The children are asleep/sleeping |
| ba- robal- il- e |
| AgrSubj2- sleep- PerfSuf-VEnd |
The perfective form -robets- ‘asleep/sleeping’ is a single morph in which the root and perfect suffix are fused together so that the boundary between root and suffix is not visible anymore.
b) Multifunctional morphs
In certain situations, it may be difficult to assign a specific meaning to a morph as it may express two or more meanings simultaneously, it may be multifunctional. Compare the subject agreement morpheme in the following:
| Morutabana o thusa bana. |
| The teacher helps the children. |
| o- thus-a |
| AgrSubj1- help-VEnd |
The subject agreement morpheme -o- expresses class (class 1), number (singular) and person (third) simultaneously.
Polysynthetic characteristics of Setswana
The term polysynthetic defines languages which are characterised by fairly long words which are made up of a large number of morphemes. Kosch (2006: 137) defines polysynthetic languages as: “A polysynthetic language is a language with a particularly high concentration of obligatorily bound morphs which bear a high semantic load”
. The most salient feature of polysynthetic languages applicable to Setswana is that there can be an abundance of morphemes per word. For verbs it is important to note that the orthography whereby verbal prefixes are written as separate orthographic word could mislead one to not recognise them as part of the verb. If recognised as a single linguistic word verbs may include a higher concentration of morphemes. There are however compounds which include different word categories such as:
| Molomatsebe |
| secret informer |
| mo- lom- a (ne-)-tsebe |
| AgrSubj1- bite VEnd NPre9-ear |
| Mojaboswa |
| heir |
| mo- j- a bo- swa |
| AgrSubj1- eat- VEnd NPre14-inheritance |
References
Berg, A. (2018) Computational syntactic analysis of Setswana. Doctoral thesis. North-West University, South Africa.
Chebanne, A. M., Creissels, D. & Nkhwa, H. W. (1997) Tonal morphology of the Setswana verb. München: Lincom.
Cole, D. T. (1955) An introduction to Tswana grammar. Cape Town: Longman.
Cole, D. T. (1975) Introduction to Tswana grammar, 2nd impression. Cape Town: Longman.
Grefenstette G. (1999) ‘Tokenization’. In Van Halteren, H. (ed.). Syntactic Word-class Tagging. Dordrecht: Kluwer. pp. 117–133.
Guthrie, M. (1967 – 1971) Comparative Bantu. 4 Volumes. Farnborough: Gregg International publishers.
Katamba, F. (1993) Morphology. London: Macmillan.
Kosch, I. M. (2006) Topics in morphology in the African language context. Pretoria: University of South Africa.
Khoali, M. H. E. (1994) Interrogative Structures in Setswana: A Functional Approach. Doctoral thesis. North-West University (University of Christian Higher Education), Potchefstroom.
Krüger, C. J. H. (1994) ‘Notes on morphology with special reference to Tswana’. South African Journal of African languages. 14(1): 15-23.
Krüger, C. J. H. (2006) Introduction to the morphology of Setswana. München: Lincom.
Krüger, C. J. H. (2013) Setswana syntax: a survey of word group structures, Volume 2. München: Lincom.
Lombard, D., Van Wyk, E. B. & Mokgokong, P. C. (1985) Introduction to the grammar of Northern Sotho. Pretoria: Van Schaik.
Louwrens. L. J. (1994) Dictionary of Northern Sotho grammatical terms. Pretoria: Via Afrika.
Matthews, P. H. (1989) Morphology: an introduction to the theory of word structure. Cambridge: Cambridge University Press.
Posthumus, L. C. (1994) ‘Word-based versus root-based morphology in African Languages’. South African Journal for African Languages. 14(1): 28–36.
Poulos, G. & Louwrens, L. J. (1994) A linguistic analysis of Northern Sotho. Pretoria: Via Afrika.
Pretorius, L., Viljoen, B., Pretorius, R. & Berg, A. (2008) ‘Towards a computational morphological analysis of Setswana compounds’. Literator. 29(1): 1–20.
Pretorius, R. S. (2014) ‘The sequence and productivity of Setswana verbal suffixes’. Stellenbosch Papers in Linguistics Plus. 44: 1–23.
Pretorius, W. J. (2000) `Die identifisering en beskrywing van die begrippe stam en wortel in die Afrikatale, met besondere verwysing na die Sothotale‘. Journal for Language Teaching. 34(1): 51–61.
Spencer, A. (1991) Morphological theory: An introduction to word structure in generative grammar. London: Wiley Blackwell.
Van Halteren, H. & Voutilainen, A. (1999) ‘Automatic taggers: an introduction’. In Van Halteren, H. (ed.). Syntactic Word-class Tagging. Dordrecht: Kluwer. pp. 109–115.
Van Rooy, B. & Pretorius, R. (2003) ‘A word-class tagset for Setswana’. Southern African linguistics and applied languages. 21(4): 203-222.
Van Wyk, E. B. (1966) ‘The word classes of Northern Sotho’. Lingua. 17(2): 230–261.
Ziervogel, D., Louw, J. A. & Taljard, P. C. (1985) A handbook of the Zulu language. Pretoria: Van Schaik.

