Extension: Character LLM¶
On this page we will learn
- how to run a local large language model (LLM) with Ollama
- how to give each character its own personality with a custom model
- how to call the model from our
Characterclass - how to handle errors when a character doesn't have a model
Terminology
- large language model – an AI model, also called an LLM, that has learnt from huge amounts of text and can write replies such as a character's dialogue.
- Ollama – a program that runs large language models on our own computer, without needing an account or internet connection once a model is downloaded.
- model parameter – one of the values a language model has learnt, used to measure its size, so a 4b model has 4 billion parameters.
- server – a program running on a computer that receives requests, such as our game's messages, and sends back replies.
- system prompt – the instructions that tell a language model who it is, which gives each character its own personality.
- Modelfile – a text file that tells Ollama how to build a custom model, including its base model, system prompt and settings.
- temperature – a model setting from 0.0 to 2.0 that controls how creative the replies are, with higher values giving more creative replies.
- terminal – a window where we type commands for the computer itself, such as
ollama, rather than Python code. - library – a collection of ready-made code that we install and import so our program can use it, such as
ollama. - try block – code inside
trythat Python runs while watching for errors, so a matchingexceptcan catch an error instead of the program crashing.
So far, our characters say the same line every time we talk to them. We can make the game more dynamic by using a local large language model (LLM) to create our characters' dialogue.
Set-up¶
We'll use Ollama to run the LLM on our own computer.
- Download and install Ollama.
- Why: Ollama runs language models locally, so no account or internet connection is needed once a model is downloaded.
- Expected result: Ollama opens with a chat box.
- Choose the gemma3:4b model from the model list in the chat box.
- Type "Hello" in the chat box.
- Why: sending a first message downloads the model.
- Expected result: after the download finishes, the model replies.
Choosing other models
We can choose other models, and this process will be mostly the same. But the bigger the model, the more memory it needs and the slower it is to reply.
Models are measured by their number of parameters: a 4b model has 4 billion parameters, while a 7b model has 7 billion. We can use Ollama's chat box to try different models and see which one works best for our game.
The list of available models is on the Ollama website.
Ollama works on two levels. There's a chat window we can type in, but it's powered by a server running on our computer. Our game can send requests to this server and get replies from the model.
Create custom models¶
Rather than using the same model for every character, we'll create a custom model for each one. This gives each character a unique personality and style. All the models use the same base model, but each has a different system prompt: instructions that tell the model who it is.
The Modelfile¶
Let's start with a model for Nigel. Create a new file in Thonny, add the text below, and save it as nigel.txt in the deepest_dungeon folder.
FROM gemma3:4b
SYSTEM You are Nigel, a friendly, but grumpy dwarf who specialises in alchemy
PARAMETER temperature 0.2
PARAMETER num_ctx 4096
Code explanation
FROM→ sets the base model for our custom model. If we chose a different model, we change this to match.SYSTEM→ sets the system prompt, which gives the character its personality. We can change this to create different characters.PARAMETER temperature→ controls how creative the replies are, from 0.0 to 2.0. Higher values give more creative replies.PARAMETER num_ctx→ controls how much text the model can take into account when it replies.
Modelfile reference
More parameters and options are in the Ollama Modelfile documentation.
Create the model¶
Now we need to open a terminal to run an Ollama command.
- In Thonny, go to Tools → Open system shell….
- Why:
ollamais a command for the computer's terminal, not for Python. - Expected result: a terminal window opens in the folder of the current file.
- Why:
- Check that the folder in the terminal prompt is deepest_dungeon.
- Type the command below and press Enter.
Code explanation
ollama→ runs the Ollama program.create→ tells Ollama to create a new model.nigel→ names the new model. We'll always name a character's model after the character, in lower case.-f nigel.txt→ tells Ollama to build the model from the file nigel.txt, which must be in the current folder.
Correct folder
If the terminal isn't in the deepest_dungeon folder, Ollama can't find nigel.txt.
Go back to Ollama's chat window. Our new nigel model should be in the model list. Select it and chat with it to get a feel for how it replies.
Adjust the model¶
To change the model, edit nigel.txt and run the ollama create command again. Try different system prompts and parameters to create different personalities.
Add the model to the game¶
Now let's connect the model to our game. We'll use a Python library called ollama to send requests to the Ollama server.
Install the ollama library¶
- In Thonny, go to Tools → Manage packages….
- Search for
ollamaand click Install.- Expected result: Thonny shows that
ollamais installed.
- Expected result: Thonny shows that
Store each character's model name¶
First, each character needs an attribute that stores the name of its model. We always name the model after the character, so Nigel's model is nigel.
Change the dunder init method in character.py as highlighted below.
Code explanation
- line 8 → creates the
modelattribute from the character's name in lower case, soNigeluses thenigelmodel.
The chat method¶
Now let's add a chat method to the Character class. It sends the player's message to the character's model and prints the reply. Add the highlighted code below.
Code explanation
- line 3 → imports the
ollamalibrary. - line 34 → defines the
chatmethod, which takes the player's message. - lines 36–39 → send the message to this character's model and store the model's reply in
response. - line 40 → prints the text of the reply.
Talk with chat¶
Now let's change the talk command in main.py so it asks the player what they want to say, and uses chat instead of talk. Make the highlighted changes below.
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 | |
PRIMM
- Predict what you think will happen when we talk to Nigel, and when we talk to Ugine. Be specific.
- Run the program. Make sure Ollama is running first.
- Time to investigate the code. What does each line do?
Code explanation
- line 69 → asks the player what they want to say to the character, and stores it in
message. - line 70 → calls the character's
chatmethod with the player's message.
Characters without a model¶
Talking to Nigel works, but talking to Ugine crashes the game with a ResponseError, because there's no ugine model. Rather than making every character have a model, let's make chat fall back to the old talk method when there isn't one.
Change the chat method in character.py as highlighted below.
PRIMM
- Predict what you think will happen when we talk to Nigel, and when we talk to Ugine. Be specific.
- Run the program.
- Time to investigate the code. What does each line do?
Code explanation
- line 36 → starts a
tryblock, which runs the code inside it and watches for errors. - lines 37–40 → send the message to the character's model, as before.
- line 41 → prints the reply, starting with the character's name.
- line 42 → catches the
ResponseErrorthat Ollama raises when the model doesn't exist… - line 43 → …and runs the character's normal
talkmethod instead.
Ollama must be running
If Ollama isn't running, chat raises a ConnectionError and the game stops. Start Ollama before playing.
Up to you¶
Now it's our turn. Can you:
- create models for your other characters, each with its own personality?
- keep a conversation history for each character, so they remember what the player said earlier?