Why Build a Language? What Creating a Conlang Taught Us About Grammar, Culture and AI

There is something wonderfully unnecessary about inventing a language.

Humanity already has thousands of them. Some are spoken by hundreds of millions of people; others survive within relatively small communities. For example, J.R.R. Tolkien, who I consider as one of my idols in this field, developed languages for the different ‘races’ that inhabited Middle Earth in The Lord of the Rings. As he wrote in one of his letters, “The invention of languages is the foundation. The stories were made rather to provide a world for the languages than the reverse. To me a name comes first and the story follows.”[1]  Taken together, conlangs display an extraordinary variety of sounds, grammatical systems and ways of organizing ideas. Some even come with fictional histories stretching back centuries or millennia.

These deliberately created languages are usually called constructed languages, or conlangs for short. Some are created for fictional worlds. Others are designed as experiments in logic or communication. Some begin as artistic projects. Some are intellectual puzzles. Some are designed with the goal of facilitating international communication. And some, some are simply made because their creators enjoy languages and want to understand them from the inside.

Now, that last reason is particularly interesting, as this was where I landed when trying to understand complex grammatical concepts in a way that would make sense to me. Because I decided, maybe foolishly in hindsight, that perhaps one of the best ways to discover how robust my actual understanding of the technical grammar really was, was to try using my understanding (along with all its faults) to build a grammar system from scratch and apply what I may (or may not!) have learned. And I can assure you, I realized how much of a novice I was, but also how this could unlock a more analytical route to unpicking the complexities of how the experts describe all languages, be they conlangs or living ones…

Taking grammar apart to see how it works

So here was my starting point. It is easy to learn a grammatical rule and believe that we understand it. I would often commend myself after an hour-long slog that I’d finally got to grips with a certain linguistic concept (only to be later disappointed that it was a tip of a much bigger iceberg – I am sure you can relate!)

Let’s take cases, for example. These can be explained reasonably quickly. As you’ll know from your own studies, a language may change the form of a noun, article or pronoun depending on the grammatical job it performs. German, for instance, retains its four-case system. Latin famously uses cases extensively and language enthusiasts will tell you, after mastering the intricacies of German, that you ‘ain’t seen nothing yet.’ But also, let’s not forget, many other languages organize the expression of similar relationships in different ways.

So… learning the rule is one thing. However, what you’ll find is, trying to design a functioning case system is something else entirely.

Suddenly you have to make decisions. Here’s what suddenly crops up:

Where is the case marked? On the noun? On an article? On an adjective? On several of them?

OK. Ahh but what happens now in the plural?

Now I want to express something more complex regarding the relationship of one idea to another. So now, I will use a preposition – but which prepositions govern which cases and why should they do so? Is it arbitrary or is there a pattern?

And then… verbs. The quagmire that is verbs! New questions emerge:

Do some verbs require particular cases or are they associated with prepositions instead? Because this is a structural decision and the answer is less clear than you’d think.

Even with nouns, which, on paper, simply have case markers, questions arise that a working language will have answered. For example, what happens when several words form a single noun phrase? Where do we multiply the rules, or where would they naturally be dropped? If we apply pure mathematics, we move away from efficient communication, so decisions are made – or in the case of working languages, conventions have prevailed, and they have done so for specific reasons.

And, perhaps most importantly, do the rules you have designed continue working when people try to say things you did not anticipate?

All of this is to say, a concept that seemed straightforward when described in a grammar book becomes a system of interacting choices when applied to the real working world of a language.

That makes conlang development an unusually powerful form of active learning. Instead of merely observing how linguistic structures work, you are forced to construct them, test them and discover where your understanding was incomplete.

And so, when I put this to the team at Lexicogs, our own experiment began in almost exactly this way. In creating Lexicogs as a really useful learning tool, one has to confirm, can we really explain concepts properly? When do we explain, when do we let convention do the talking, what does someone need to know and at what point in their language journey – and from that, how do we create lessons that fulfil this need. And this varies from one language to another. So we decided, the only way to really see this clearly was to put ourselves to the test, and so the experiment of developing a conlang was born.

So, to be clear, the original intention was not to create an elaborate fictional civilisation (more on this later). It was much more modest: to explore grammatical systems, starting with case and verb systems, by trying to build one.

The next question was: where to start? We decided to look specifically at Germanic languages and use that as the inspiration, not least because English itself is a Germanic language, albeit with lots of new layers on top, but this seemed like a good starting point.

The vocabulary therefore was largely developed from Proto-Germanic and then the wider Germanic language family. In practice, the grammar gradually became more sophisticated as we explored examples. Words accumulated. Usage examples were created.

And then something unexpected happened.

When vocabulary starts suggesting a culture

Not part of the original plan, but what became increasingly clear, was that languages are not simply lists of interchangeable labels – we all know this as language enthusiasts, but what we began to see was the mechanics of what drives this.

As you’ll have no doubt experienced, an English word and its apparent equivalent in another language rarely occupy precisely the same semantic territory. Languages divide experience differently. And the same is exactly, and perhaps surprisingly true, of a conlang.

Experience shows us, one language may distinguish concepts that another expresses with a single word. Another may use the same word across ideas that English treats separately. Metaphors become conventional. Historical meanings linger inside apparently ordinary expressions.

But… When building a language carefully, exactly the same thing begins to happen.

At first, you consciously choose meanings.

Eventually, relationships between those meanings begin suggesting things you did not consciously plan.

A vocabulary concerning landscape may suggest how the speakers relate to their environment. Words surrounding obligation, protection, hospitality or judgement may begin implying assumptions about society. Distinctions within apparently simple verbs can reveal what kinds of actions the language treats as meaningfully different. These all become not just word choices, but carefully considered choices.

At that point something curious happens.

You are no longer simply asking:

“What word should this language have?”

You begin asking:

“What kind of people would need these distinctions?”

And that question opens an entirely different door.

Geography begins to appear. Social structures develop. Stories become possible. Moral assumptions emerge. Customs require explanation that you didn’t expect.

Suddenly, the language is no longer sitting inside an experimental framework, but a conlang is developing a fictional world of its own.

In other words, the fictional world is beginning to grow out of the language.

Language as a way of exploring worldview

This does not mean that grammar magically determines culture. That would be a much stronger claim, and one we would not make – and there are academics who indeed will be devoting their life’s work to exploring this idea, so it is not our place to say.

What we can, however, say is that what language appears to do is create conceptual possibilities. This conclusion was much stronger than we ever thought it would be.

To explain, we found:

Once a language distinguishes two ideas, those ideas become easier to discuss separately.

Once a metaphor becomes productive, it can be extended.

Once a society has vocabulary for a particular kind of obligation, judgement or relationship, writers using that language naturally begin exploring what that concept implies.

This is one reason conlangs can become unexpectedly absorbing.

They are simultaneously creative and analytical.

You design a system – and then start discovering unintended consequences inside the system you designed.

That can lead somewhere considerably deeper than naming imaginary mountains. Rather, it can become a form of a much larger cultural thought experiment.

What would authority mean in a society that linguistically associates power with protection?

What would justice look like if restitution were conceptually distinct from punishment based on a word choice we have decided upon?

What happens if myths are classified not according to whether they are literally factual, but according to the kind of truth a community believes they preserve?

What sort of stories would such a culture tell?

At that point, vocabulary, literature and philosophy begin reinforcing one another.

And the conlang becomes less like a code and more like an ecosystem.

Then AI entered the experiment

There is another reason this particular project developed in an unusual direction.

We were also exploring what artificial intelligence could –  and could not – contribute. This is also important when we assess the role of AI in developing language courses for working languages.

Now, at first the possibilities appeared obvious:

Large language models are extraordinarily good at recognising linguistic patterns. They can analyse grammar, propose examples, identify relationships and produce large amounts of text extremely quickly.

Surely, then, a constructed language would be an ideal environment for AI.

The reality was considerably messier. Here’s why:

AI could produce extremely convincing mistakes.

Give a model a language with Germanic ancestry and it could happily invent a word that looked magnificently Germanic, obeyed the apparent sound system and fitted the sentence beautifully.

There was only one problem.

The word did not exist.

The same problem appeared with grammar. Once the model recognised a productive pattern, it might apply that pattern somewhere the language had never authorised it.

The result was especially dangerous because incorrect output often looked completely plausible.

This led to one of the most useful lessons of the entire experiment:

Plausibility is not the same thing as correctness.

And that matters far beyond constructed languages.

Teaching an AI what it is allowed to know

The solution emerged slowly through repeated failure.

A reliable system needed at least three different forms of authority (it may need more, but we haven’t discovered them yet!) But let’s focus on what we did find that we needed, which emerged through these repeated failures:

First came a controlled lexicon: the words that actually existed, together with their meanings, grammatical properties and accepted forms.

Then came an explicit grammar: the rules governing how those words could legitimately interact.

Finally came an attested corpus: examples of the language being used in conversation, narrative, songs and other contexts.

In the event, we can conclude that, despite our best wishes, none of these proved sufficient alone.

A dictionary could tell the model which words existed but not everything about how speakers actually used them.

A grammar could provide productive rules while still encouraging the model to generate forms that appeared structurally possible but had never been admitted into the language.

Examples showed real usage, but without the grammar it was difficult to distinguish an underlying rule from an accidental feature of one sentence.

Together, however, they produced something much more powerful.

Lexicon. Grammar. Corpus.

Each constrains the others.

And an important principle related to AI gradually emerged: when the model is uncertain, it must not fill the gap.

It must stop.

That sounds simple, but it reverses one of the normal expectations placed upon generative AI. We usually ask models to continue producing useful answers.

In a controlled linguistic system, sometimes the most valuable answer is:

“That concept does not yet exist in the canon.”

The goal therefore became not teaching the AI to invent more successfully, but teaching it when not to invent at all.

What failure taught us

This process took a surprisingly large number of iterations, and it comes after nine months of repeatedly getting it wrong.

Words drifted, grammar was inconsistently applied, earlier examples contradicted later rules and then needed to be revised, over and over until the system – both in terms of the AI model and the rules of the conlang itself – were stable.

Then, when we came to producing artefacts in the language, new problems with today’s AI tools emerged that we hadn’t built into the design. For example, even using phonetic spelling based on IPA, speech synthesis pronounced perfectly valid words either incorrectly or in artificially unnatural human speech patterns.

And then, music-generation systems interpreted spelling according to entirely different linguistic assumptions, despite being given what superficially seemed like entirely robust instructions.

Changes made in one part of the language unexpectedly affected another and everything was more tightly, but inextricably interwoven in ways we did not expect.

But, these were healthy failures. Indeed, those failures became part of the experiment and contributed even more to our understanding of how languages really operate.

Every problem required another distinction between what was assumed and what was actually defined. Over time, the project developed increasingly clear ideas about source authority, linguistic validation and the division of responsibility between human creator and AI collaborator.

The AI became most useful not when it was permitted unlimited creativity, but when creativity occurred inside carefully maintained boundaries.

That is perhaps the most transferable lesson we have learned so far.

The interesting part is the loop

There is also a temptation when discussing AI-assisted creative work to ask who created something.

Did the human create it or did the AI create it?

Our experience suggests that this can sometimes be the wrong question.

The interesting thing is the loop. A human establishes a linguistic rule or chooses a word – in our case, we decided what the master Germanic root word should be by looking at English, Old English, Icelandic, Norse fragments, German, Dutch and Proto-Germanic, or even Indo-European constructions when necessary. We took a view on what we liked, what made sense semantically, what made sense culturally and what fitted best a human preference.

Following this, the AI leans into its strongest abilities wonderfully. It identifies the possible implications, the patterns, the wider relationships and links, and, most importantly, the contradictions — because holding that whole mental map of attested and non-attested relationships simultaneously becomes increasingly difficult as the system grows.

So, the AI lands its analysis. The human judges whether those implications are linguistically and culturally convincing and makes a decision whether this forms part of an official attested usage by testing the idea through actual use.

And, in fact, the contradictions are the most useful part. Based on these, the language is revised and made into a more coherent system. The revised attested system then makes new ideas possible. Importantly, neither participant simply dictates the finished result. The productive unit is the iteration and the lessons this provides and the enlightenment that follows.

So, to turn back to our original experiment: it has turned a relatively small grammatical exploration exercise into something we did not originally expect and could not at all have anticipated: a new conlang, a language that now has thousands of controlled lexical entries, and with that, its own developing literary forms, social concepts, mythology, music, narratives and a fictional cultural world. Being language nerds, this has become an almost ‘as important’ side project to our original goal of providing really immersive language lessons.

For us, it has also raised an intriguing question, that we did not expect to be debating in the field of conlanging. And that is: When does a constructed language stop being merely an invented communication system and start becoming a way of thinking through an imagined culture?

We do not yet know the answer.

But exploring that question has become part of this unintentional ‘side-project.’ We’re now looking at this in the broader context of AI, its uses in language learning, and the new artistic creations it can spawn, as we’d love to be part of the future discussion on this subject matter and contribute where we can.

So why build a language?

OK, so through our own experiences, we feel that we may be better placed to answer this than we were at the beginning, although we do not claim to be in any way the experts in this emerging field of linguistics, far from it!

And the annoying answer, but the correct one seems to remain: there are as many answers to this question as there are people making them.

Some want to populate fictional worlds – and in the creative arts space this will continue to be hugely valuable, and we can’t wait to see the influence and potential of AI here.

 Some love phonology – and AI creates an amazing opportunity to explore how ‘dictated rules’ morph when confronted with human conversational reality.

Some enjoy grammatical systems. There’s an interesting question that emerges here. Can a language be mathematically perfect? Should it be? It enlightens us as to why ‘exceptions’ exist in real languages when grammar rules hit the reality of efficient human communication.

Some want to experiment with communication. There’s a strong case, culturally, for developing a language that is ‘politically neutral,’ but is it possible within our global culture?

Some simply find pleasure in discovering what happens when sounds, grammar and meaning are assembled differently. We think that this, combined with the possibilities of AI, could open fascinating new avenues for linguistic experimentation: helping us explore how linguistic systems emerge, why particular kinds of change occur, and what those experiments might suggest about the continuing evolution of living languages.

And sometimes, just sometimes, a project begins for one reason and quietly becomes something else.

That is what happened to ours.

What started as an attempt to understand language structure gradually became an experiment in linguistic design, cultural emergence and AI-assisted creativity.

We call the language Boersk.

And somewhere behind it now sits an entire fantastical world that was never part of the original plan.

If you are curious, you can explore the Boersk project here.

In conclusion

Even if invented languages themselves hold no particular appeal, and there’s no reason they should, the experiment has already reinforced something central to the philosophy behind Lexicogs and what we already believed:

Languages are not simply collections of words to memorise.

They are systems through which people organise experience, preserve culture, express relationships and make meaning. And for those with a curious appetite, we’d say this… Sometimes the best way to appreciate how extraordinary those systems are is to try building one and just seeing where that journey takes you!


[1] J.R.R. Tolkien, Quoted in Humphrey Carpenter and Christopher Tolkien (eds.), No. 165, The Letters of J.R.R. Tolkien

Leave a Comment

Your email address will not be published. Required fields are marked *

Are you human? Please solve:Captcha