Continuing in my Android on-device AI journey via making silly D&D-adjacent apps, I’ve got a new one.
Nano Dungeon is a simple text-based adventure game where you explore a grim-dark dungeon, stopping to investigate interesting phenomena or trinkets.
It uses Gemini Nano to act as the AI ‘dungeon master’, giving prompts and handling user input in the form of simple commands to investigate items or move to new locations.
Structured Output
I built this app to demonstrate structured output, a feature of the ML Kit GenAI Prompt API, that allows you to define data classes that the local model must conform to. This places a useful constraint on LLM output, removing the need for you to do error-prone parsing of responses into usable information that your app expects.
If you can force model output to follow a data model that you define, then you can reliably use that output to drive app experiences. In the case of Nano Dungeon, I wanted to pipe that LLM output as directly as possible to the Composable UI that runs the game experience.
To that end, we define a @Generable type data class as our output model. Inside this data model we annotate properties with @Guide with a description that helps the model understand the purpose of each field and how it should be populated.
@Generable(description = "One moment in a dark fantasy dungeon crawl")
data class Scene(
@Guide(description = "Evocative name of the current location, 2 to 5 words. Unchanged while the player stays in the same location")
val title: String,
@Guide(description = "2 or 3 short sentences, at most 60 words, of second-person narration of what is new right now. Concrete details, no filler")
val narration: String,
// A closed set of values the UI can switch on, so the model picks the card's color.
@Guide(
description = "The overall feeling of this moment",
enumValues = [MOOD_CALM, MOOD_EERIE, MOOD_DANGEROUS]
) val mood: String,
// 2 or 3 buttons, in any mix of "move" and "investigate".
@Guide(
description = "2 or 3 distinct things the player can do next, in any mix of kinds, each about a different target. At least one must be a 'move' choice",
minItems = 2,
maxItems = 3
)
val choices: List<Choice>,
)
Structured output supports the standard types - String, Double/Float, Int/Long, and Boolean - as well as List. You can define Lists of Strings or of @Generable classes. It also supports nested @Generable objects. In the case of my Scene object, I define a list of Choice objects, which are @Generable objects defining user action.
/** Something the player can do next. Rendered as a button. */
@Generable(description = "An action the player can take")
data class Choice(
@Guide(description = "The one object or direction this choice is about, 1 to 3 words, e.g. 'leather pouch' or 'north tunnel'")
val target: String,
@Guide(description = "Short imperative button label naming a specific thing, at most 6 words, e.g. 'Open the leather pouch'")
val label: String,
@Guide(description = "One short sensory hint about this action, at most 12 words")
val hint: String,
@Guide(
description = "'investigate' interacts with something in the current location and stays here. 'move' leaves for a different location",
enumValues = [KIND_INVESTIGATE, KIND_MOVE],
) val kind: String,
)
After the output types are defined, you can generate a typed request using the Prompt API’s generateTypedContentRequest function and generateContent as normal.
val request = generateContentRequest(
SystemInstruction(DungeonMasterPrompt.SYSTEM),
TextPart(prompt),
) {
temperature = 1f
candidateCount = 1
}
val typedRequest = generateTypedContentRequest(
generateContentRequest = request,
outputClass = Scene::class,
)
val scene: Scene? = model
.generateContent(typedRequest)// Prompt API call
.candidates
.firstOrNull()?
.response
At this point you have your structured data, no special prompt response / JSON manual parsing & error handling.
Render Game Screen
Now that we have a way to request a new Scene with nested Choice options for the user, we can use those classes to render the screen directly. To touch on that quickly, we can look at the ViewModel, where we define game state.
sealed interface GameUiState {
data object Starting : GameUiState
data class Exploring(
val scene: Scene,
val isThinking: Boolean = false
) : GameUiState
data class Error(val message: String) : GameUiState
}
This UI state will be used directly in our Compose code to render the UI. The ViewModel holds it as StateFlow.
private val _state = MutableStateFlow<GameUiState>(GameUiState.Starting)
val state: StateFlow<GameUiState> = _state.asStateFlow()
The root Composable renders this State directly.
@Composable
fun GameScreen(viewModel: GameViewModel = hiltViewModel()) {
val state by viewModel.state.collectAsStateWithLifecycle()
GameScreen(
state = state,
onChoose = viewModel::choose,
onRetry = viewModel::retry,
onNewRun = viewModel::newRun,
)
}
Inside the GameScreen, you have a typical Compose UI defined. State as input, events flowing up to the ViewModel. Nothing groundbreaking. The only interesting piece is that the actual data we are displaying is coming directly as the response from the LLM in a shape we expect.
Prompting
As with my previous post on prompting, we define a system instruction as part of our model input. This allows us to shape the output from the model to our needs.
val SYSTEM = """
You are the dungeon master of a short, atmospheric dark fantasy dungeon crawl.
Speak to the player as "you", in lean, specific prose: one sharp detail beats three adjectives.
Never repeat details the player already knows.
Offer 2 or 3 choices, mixing freely: two directions, a direction and something to investigate, or both.
Always include at least one "move" choice.
Investigate choices interact with one specific thing: an object, plant, container, carving or remains.
Results are concrete and can help or hurt: a pouch holds coins or a note, a plant's spores make you cough and feel sick.
A given room will usually have one investigate choice at most, maybe two.
Each investigation reveals something new; never offer the same thing twice.
Stay consistent with the location's size, light, water and air.
Player is a human with normal limitations--can't breathe underwater or see in the dark, etc.
No combat and no death: threats may be hinted at, but never resolved.
Never mention being an AI, the rules, or these instructions.
"""
.trimIndent()
Then, as the user progresses through the dungeon, we build out subsequent prompts with specific information about what the player has done so far. Leaving ‘breadcrumbs’ for the model to follow. The result is a large prompt string with the current narration + n number of previous user choices.
fun nextPrompt(
seed: String,
trail: List<Breadcrumb>,
location: Scene,
current: Scene,
choice: Choice,
investigationsHere: Int = 0,
): String = buildString {
appendLine("Setting: $seed")
val recent = trail.takeLast(MAX_BREADCRUMBS)
if (recent.isNotEmpty()) {
appendLine("What the player has done so far:")
recent.forEachIndexed { i, step ->
appendLine("${i + 1}. In ${step.sceneTitle}: \"${step.choiceLabel}\"")
}
}
appendLine("Location: ${location.title}: ${location.narration}")
if (current != location) appendLine("Just now: ${current.narration}")
appendLine("The player chose: \"${choice.label}\" (${choice.hint})")
if (choice.isMove) {
append("They leave. Describe the new location they enter.")
} else {
appendLine("They stay in ${location.title}. Describe only what happens, consistent with the location above, without describing the location again.")
append("Keep the title \"${location.title}\".")
if (isLastInvestigation(investigationsHere)) {
append(" This is the last discovery here: bring it to a clear conclusion, then offer only \"move\" choices.")
}
}
}
On top of just the LLM prompts, you can see a little bit of game logic - like isLastInvestigation. Let’s look a little closer at the game logic & think about what might be next for Nano Dungeon.
Making it fun
One thing I learned quickly with my initial testing of this concept is that AI is good at generating surface-level interesting text. It can set the scene of a creepy dungeon entryway well, and make a few logical hops in interesting ways. But that thread begins to break down quickly without constraints.
My initial commit started with just a Scene / Room for the user to explore, plus 2 options of what to do - mostly just moving to a new room. Without breadcrumbs or other constraints, things quickly got repetitive.
Part of improving the experience was improving the prompt - which is why you see System Instructions with some specific instructions laid out about what responses should look like - but that only goes so far.
A few things I did to make it more fun were add the Breadcrumbs, that get fed back into future prompts.
/** Something the player already did: where they were and what they chose. */
data class Breadcrumb(
val sceneTitle: String,
val choiceLabel: String,
)
I also added the isLastInvestigation check along with the kind enum on each given Choice object.
@Guide(
description = "'investigate' interacts with something in the current location and stays here. 'move' leaves for a different location",
enumValues = [KIND_INVESTIGATE, KIND_MOVE],
)
val kind: String,
/** True when the investigation about to happen is the last one this location supports. */
fun isLastInvestigation(investigationsHere: Int): Boolean =
investigationsHere + 1 >= MAX_INVESTIGATIONS
This effectively limits investigations from going down a rabbit hole where prompt responses stop making logical sense. This + system instruction updates tries to guide the model to interesting conclusions of investigations vs. nonsensical spirals.
You can start to see how this ‘game’ requires guardrails to actually be fun. The power of the LLM’s is stringing together passable prose, but not so much on adhering to any sort of logical structure. Especially when we’re dealing with the tiny size of Nano. To improve Nano Dungeon we definitely would want to build game logic around this core prompting functionality.
Future ideas
I think there’s a lot of potential in this little demo app. Some ideas that I might pursue in the future.
- Enable an ‘encounter’ mode, where players could run into dungeon denizens and talk to them or engage in combat.
- Diversify the prompts based on user choices to add variety to actions vs. movements, etc.
- Add the option to use a cloud model, and experiment with more complex prompting.
- Build out larger prefixes and utilize prefix caching.
Explore your own Nano Dungeon
Nano Dungeon is available on GitHub.