@inbrowser/model
The model layer for the stack. It owns the one model-call contract —
ModelClient — plus the cloud providers that implement it and the
on-device LLM engine. @inbrowser/relay (transport) and
@inbrowser/agent (runtime) both consume a ModelClient, so this is the
single shared definition of “an LLM” for everything downstream.
Three public seams, one package:
- Root (
@inbrowser/model). DefinesModelClient/ModelRequest/ModelEvent, plus usage helpers andwithRetry. No provider factories. - Providers (
@inbrowser/model/providers/<name>). Each cloud provider (gemini,openrouter,requesty,anthropic,oai-compat,ollama,llama-server,claude-cli,claude-code) and the Firebase AI Logic constructed-model adapter (firebase-ai-logic) returns aModelClient. - Local (
@inbrowser/model/local).createEngineloads ONNX models in the browser via@huggingface/transformers+ ONNX Runtime Web (WebGPU / WASM) behind a narrowEnginesurface that streamsEngineEvents.
Status. Contract + cloud providers are the live integration path: relay and agent both consume a
ModelClient.createEngineloads a model through@huggingface/transformersandgenerate()streams real tokens (the end-to-end load path runs inexamples/local-llm-poc, headless-verified). The engine is now aModelClienttoo, viacreateEngineModelClient(from@inbrowser/model/local), which widens the engine’sEngineEventstream to the contract’sModelEvent. The old@inbrowser/model/relayand@inbrowser/model/agentadapter subpaths have been removed. Known gaps:GenerateOpts.stopsequences are accepted but not yet enforced, and the site’s in-browser docs-chat path that drives a local engine through the agent is still forthcoming (the adapter exists; the site toggle does not).
A cloud model as a ModelClient
import { geminiModelClient } from '@inbrowser/model/providers/gemini';
const client = geminiModelClient({ apiKey: process.env.GEMINI_KEY, model: 'gemini-3.5-flash' });
for await (const evt of client.chat(
{
messages: [{ role: 'user', text: 'Explain WebGPU in one paragraph.' }],
tools: [],
toolUseEnabled: false,
},
new AbortController().signal,
)) {
if (evt.kind === 'text') process.stdout.write(evt.text);
else if (evt.kind === 'usage') console.error(evt.usage);
}The turn ends when the iterable returns; a usage event (or a terminal
error event) is the last thing emitted. There is no turn_complete
event.
Firebase AI Logic from an existing Firebase app
Firebase AI Logic uses the Firebase app that the host already configured. The
host owns Firebase initialization, App Check, backend selection, Vertex AI
location, and construction of the GenerativeModel; this package only adapts
that model to ModelClient:
import { initializeApp } from 'firebase/app';
import { getAI, getGenerativeModel, GoogleAIBackend } from 'firebase/ai';
import { createFirebaseAiLogicModelClient } from '@inbrowser/model/providers/firebase-ai-logic';
const app = initializeApp(firebaseConfig);
// Initialize App Check for `app` before making production AI requests.
const ai = getAI(app, { backend: new GoogleAIBackend() });
const firebaseModel = getGenerativeModel(ai, { model: 'gemini-3.5-flash' });
const client = createFirebaseAiLogicModelClient(firebaseModel);The adapter streams text and thinking, translates caller-run custom function
calls (including thought-signature replay), maps sampling and usage, forwards
cancellation, and normalizes Firebase errors. It has no firebase dependency:
the constructed model crosses a small structural interface. Live API, Imagen,
server templates, Firebase built-in/automatic tools, multimodal events,
structured-output configuration, token counting, and hybrid lifecycle control
are intentionally outside this adapter. See the
Firebase AI Logic reference.
A local OpenAI-compatible server
Ollama, llama.cpp’s llama-server, vLLM, LM Studio, LocalAI, and friends all
expose the same OpenAI POST /v1/chat/completions wire shape. One generic
factory talks to any of them; two named presets carry the right defaults for
the common local servers:
import { openaiCompatModelClient } from '@inbrowser/model/providers/oai-compat'; // any OAI server
import { ollamaModelClient } from '@inbrowser/model/providers/ollama'; // localhost:11434
import { llamaServerModelClient } from '@inbrowser/model/providers/llama-server'; // localhost:8080
// Generic: point at any OAI-compatible server. `apiKey` becomes a Bearer token.
const vllm = openaiCompatModelClient({ baseUrl: 'http://gpu.local:8000', model: 'qwen2.5' });
// llama.cpp llama-server. `--api-key` is optional; pass it as `apiKey`.
const llama = llamaServerModelClient({ model: 'qwen2.5-coder', apiKey: process.env.LLAMA_KEY });Tool calling on
llama-serverneeds--jinja. The server only honors the OpenAItoolsarray when launched with--jinja(so it applies a tool-aware chat template); without it, tool calls never stream back. Auth is off unless you start it with--api-key KEY.
The presets delegate to openaiCompatModelClient; reach for the generic factory
directly for any server without a named preset.
An on-device model via the engine
The local engine uses the optional Transformers peer. Install it only in an application that runs on-device inference:
npm install @inbrowser/model @huggingface/transformersThen import the opt-in local surface:
import { createEngine, gemma4_E2B } from '@inbrowser/model/local';
const engine = createEngine(gemma4_E2B);
await engine.ensureReady();
for await (const evt of engine.generate([
{ role: 'user', text: 'Explain WebGPU in one paragraph.' },
])) {
if (evt.kind === 'token') process.stdout.write(evt.text);
}The engine speaks EngineEvent (token / thinking / tool_call /
usage / error), not ModelEvent. To use it as a ModelClient —
e.g. to hand it to the agent — wrap it with createEngineModelClient:
import {
createEngine,
createEngineModelClient,
smollm2_360m,
} from '@inbrowser/model/local';
const engine = createEngine(smollm2_360m);
const client = createEngineModelClient(engine); // a ModelClient
for await (const evt of client.chat(
{ messages: [{ role: 'user', text: 'Hello' }], tools: [], toolUseEnabled: false },
new AbortController().signal,
)) {
if (evt.kind === 'text') process.stdout.write(evt.text);
}The adapter maps token → text, folds the engine’s terminal usage
into a ModelEvent usage, passes tool_calls through (no signature),
and drops the engine-only extras (decodeMs, recoverable). Wiring a
local model into the docs-chat site through the agent is forthcoming;
the createEngineModelClient building block it needs now exists.
Surfaces
The root @inbrowser/model is the lightweight contract/provider surface:
| Export | What it gives you |
|---|---|
ModelClient, ModelRequest, ModelEvent, ModelMessage, ModelUsage, ToolSpec, ReasoningEffort | The shared contract (type-only) |
geminiModelClient, openrouterModelClient, requestyModelClient, anthropicModelClient, openaiCompatModelClient, ollamaModelClient, llamaServerModelClient, claudeCliModelClient, claudeCodeModelClient | Cloud + local provider factories; each returns a ModelClient |
createFirebaseAiLogicModelClient(model, opts?) | Wraps a caller-constructed Firebase AI Logic GenerativeModel; Firebase/App Check lifecycle stays with the host |
OpenAiCompatConfig, OllamaConfig, LlamaServerConfig | Config shapes for the OpenAI-compatible factory and its local presets |
withRetry(client, opts?) | Decorator that retries transient upstream errors while nothing has streamed |
CloudProviderConfig, ModelClientFactory | Shared provider config + the factory type the relay routes on |
The opt-in @inbrowser/model/local surface contains everything tied to
on-device inference:
| Export | What it gives you |
|---|---|
createEngine(preset) | Runtime Engine — owns load state + decode loop, streams EngineEvent |
createEngineModelClient(engine, id?) | Wraps an Engine as a ModelClient (maps EngineEvent → ModelEvent) |
definePreset(p) | Type-safe identity helper for community presets |
parseToolCalls, splitThinking | Stream transformers over an EngineEvent stream |
ModelPreset, Engine, EngineEvent, … | Public engine types |
gemma4_E2B, gemma4_E4B, qwen2_5_coder_1_5b, qwen3_1_7b, deepseek_r1_qwen_1_5b, smollm2_360m | The bundled presets |
hostEngineInWorker(self), connectWorkerEngine(opts) | Worker host/connect helpers |
Consumers that do not import /local never cross the on-device runtime seam.
Vocabulary anchor
- ONNX — model file format. ONNX Runtime Web is the execution
engine (
onnxruntime-web); WebGPU and WASM are its backends. dtype— weight/activation precision selection (q4f16,q8,fp16,fp32). Distinct from parameter count.ModelRef— bare locator (HF HubmodelId+revision).ModelPreset— locator + dtype + backend + capabilities. Static.Engine— runtime object owning a loaded model. Dynamic.- Cold start — fetch + init + warmup. Warm decode — subsequent calls on a ready engine.
Design notes
- One factory (
createEngine), many presets. NocreateGemmaEngine. capabilitiesis on the preset, not the engine — interrogable pre-load (gemma4_E2B.capabilities.contextWindow).EngineEventis narrower than the contract’sModelEvent(no cost, nothoughtSignature).createEngineModelClientis the place that widens it — translate at that boundary, not in the engine.- Worker subpath returns the same
Engineshape; a consumer cannot tell whether it holds a direct or remote engine. - Tool calling is not native to Gemma 4. The polyfill (prompt-engineered
tool calling + structured-output parsing) lives in
@inbrowser/agent, not here.