Home Articles Resume Nala Project
ES EN

Nala AI Runtime Architecture: Genkit, Tool Orchestration and Safety-by-Design

Nala AI Runtime Architecture: Genkit, Tool Orchestration and Safety-by-Design

Nala AI Runtime Architecture: Genkit, Tool Orchestration and Safety-by-Design

Nala is not designed as a simple call to a model. The interesting part is the architecture around the LLM: a dedicated runtime that prepares context, applies safety, controls tools, executes Genkit, validates output, updates memory and records traceability.

This matters because an AI demo can survive with a large prompt and a direct SDK call. A real product cannot. Once memory, family rules, streaming, parent signals, tools and safety appear, clear boundaries become necessary.

Nala AI Runtime Architecture

1. The architecture problem

A first AI integration often starts like this:

const response = await openai.chat.completions.create(...)

That is enough to validate an idea, but it does not answer product questions:

  • where conversation context lives;
  • how family rules are applied;
  • how sensitive messages are handled;
  • which tools the model can use;
  • how to avoid tools using the wrong adapter;
  • how to persist memory without storing dangerous data;
  • how to debug a bad response;
  • how to stream without duplicating all logic.

Nala separates those concerns into layers. The API adapts HTTP. The runtime governs execution. Genkit orchestrates prompts and the model. Adapters encapsulate persistence and context. Policies control safety, memory and tools.

2. API boundary

The main endpoints are:

api/nala/chat.ts
api/nala/stream.ts

Their responsibility is deliberately small: receive the request, validate the public contract, build the internal input and delegate to the runtime.

The normal endpoint ends in:

defaultNalaRuntime.runChatFlow(...)

The streaming endpoint ends in:

defaultNalaRuntime.runChatStreamFlow(...)

Delivery changes, but the execution model remains the same.

3. Runtime composition

The center lives in src/runtime.ts:

createNalaRuntime(deps)

This function creates a runtime with explicit dependencies:

  • adapters;
  • tools built on those adapters;
  • internal services;
  • allowLlmToolUse flag;
  • runChatFlow and runChatStreamFlow.

This avoids global-singleton coupling. An isolated runtime can use injected adapters and preserve clear guarantees in tests, separate environments or future multi-tenant configurations.

The default runtime uses inMemoryNalaContextAdapter and enables allowLlmToolUse: true for Developer UI and global flows. Custom runtimes default to allowLlmToolUse: false.

4. Real execution sequence

A turn is not “prompt → model → response”. It is a pipeline.

Runtime request sequence

The main steps are:

const trace = createTrace('nalaChatFlow', input.metadata?.traceId)
const prepared = await prepareNalaTurn(input, trace, services)
const promptInput = buildChatPromptInput(input, prepared)
const relevantTools = selectRelevantTools(input, prepared, services)
const fallback = fallbackPromptOutput(input, prepared)
const { skip, reason } = shouldSkipChatLlm(prepared)

If the LLM is not skipped, the runtime executes:

const response = await nalaChatPrompt(promptInput, {
  tools: relevantTools.actions,
  maxTurns: 3,
  returnToolRequests: true,
})

And finally:

return finalizeNalaTurn({ input, prepared, promptOutput, trace, services })

The fallback is prepared before the model call. If the LLM fails or safety suggests avoiding it, the system can still return a controlled response.

5. Genkit inside Nala

Genkit is initialized in src/genkit.ts:

export const ai = genkit({
  name: 'nala-ai-api',
  plugins: [openAI()],
  model: `openai/${resolveNalaModelName()}`,
})

In Nala, Genkit provides:

  • definePrompt;
  • input and output schemas;
  • provider abstraction;
  • streaming;
  • tool integration.

But Genkit does not own all business logic. The healthier architecture is to use Genkit as the AI execution layer while the runtime keeps product decisions.

6. Contracts-first

Contracts live in src/contracts/nala.contracts.ts.

Key schemas include:

  • NalaChatInputSchema;
  • NalaChatPromptInputSchema;
  • NalaChatPromptOutputSchema;
  • NalaChatOutputSchema;
  • NalaSafetyOutputSchema;
  • NalaIntentOutputSchema;
  • NalaSessionUpdateSchema.

This reduces ambiguity. The API does not receive arbitrary input. The prompt does not consume an improvised object. The final response is validated before reaching the client.

In AI systems this matters because model failures can be unusual. Contracts reduce the blast radius.

7. Control plane around the model

The right way to view Nala is as a control plane around the LLM.

LLM control plane

The model is surrounded by four boundaries:

  1. contracts;
  2. safety policies;
  3. memory rules;
  4. tool policy.

The LLM answers, but it does not decide by itself what context it sees, what tools it can use or what gets persisted.

8. Tool orchestration

Tool calling is one of the most delicate parts of any LLM architecture.

Nala separates two concepts:

createNalaTools(adapters)

creates internal tools closed over injected adapters.

Genkit tools are declared separately with ai.defineTool.

Tool orchestration model

Per-turn selection happens through:

selectRelevantTools(input, prepared, services)

In addition, allowLlmToolUse controls whether actions are actually passed to the model.

This avoids a subtle problem: Genkit tools are global. If they are passed into an isolated runtime carelessly, they may end up using global adapters instead of injected adapters. That is why custom runtimes do not allow LLM tool use by default.

9. Safety-by-design

Safety should not be only a line inside the prompt. In Nala it is part of execution.

prepareNalaTurn integrates intent, safety, memory and family rules before the final prompt is built. The safety flow can return:

  • safetyLevel;
  • allowed;
  • categories;
  • blockedReason;
  • redirectionStrategy;
  • parentSignal.

This allows the runtime to decide whether to call the LLM, use fallback or record a parent signal.

For a child/family assistant, that separation is not optional. Privacy, tone and boundaries are part of the product.

10. Memory strategy

Nala does not try to remember everything.

The strategy is to preserve continuity without turning memory into a dangerous container:

  • capped recent messages;
  • session summary;
  • safe preferences;
  • sanitization before saving;
  • rejection of sensitive data;
  • update at the end of the turn.

Useful memory is usually small. Remembering “she likes dinosaurs” can improve the experience. Storing sensitive personal data should not.

11. Streaming

Streaming is implemented as a variant of the same runtime.

Streaming lifecycle

Execution calls:

nalaChatPrompt.stream(promptInput, {
  tools: relevantTools.actions,
  maxTurns: 3,
  returnToolRequests: true,
})

Chunks are emitted progressively and, at the end, the complete structured response is awaited, parsed and finalized.

This avoids a second architecture for streaming. Normal and streaming paths share preparation, safety, tool selection, fallback and finalization.

12. Observability

Nala records traces with:

  • createTrace;
  • addTraceStep;
  • recordToolCall;
  • completeTrace.

This shows detected intent, safety level, selected tools, fallback usage and how the turn completed.

In AI systems, observability is not optional. Without traces, debugging bad answers becomes guesswork.

13. Tests and evals

The project includes tests and evaluation datasets for:

  • safety;
  • privacy;
  • continuity;
  • style;
  • games;
  • family rules;
  • memory;
  • tools;
  • runtime.

That is a good sign. AI applications should not be validated only by trying a few prompts manually. They need repeatable regression cases.

14. Tradeoffs

This architecture increases control and testability, but it adds more moving parts:

  • more contracts to maintain;
  • more runtime steps;
  • more observability points;
  • more discipline around tool registration;
  • more care with memory and adapters.

For a sensitive assistant, that cost is justified. The alternative —one huge prompt with too many responsibilities— scales worse.

15. Future evolution

Natural next steps would be:

  • persistent adapters with Redis or Postgres;
  • prompt and policy versioning;
  • trace export to OpenTelemetry;
  • a parent-signal review panel;
  • a retrieval layer for external knowledge;
  • automated evals with thresholds;
  • clearer separation between production runtime and Developer UI.

Conclusion

The interesting part of Nala is not only that it uses Genkit.

The interesting part is how it uses Genkit: inside a runtime that controls context, safety, memory, tools, fallback, streaming and final output.

That design is what separates a chatbot demo from a serious foundation for a production AI assistant.

#Nala#Genkit#OpenAI#TypeScript#runtime#safety#tool-calling

Alex Sanz

I build products and systems where architecture, business needs, and AI become real, reliable, and maintainable capabilities.

Related articles

View all