Adding an AI model
Explicit instructions work without a model. Connect one to understand free phrasing like “Indian customers”, “big spenders” or “people who haven't logged in for a while”. You bring the model and the key; Pragma never ships one.
Explicit instructions are answered here, in about 1 ms. No network.
Holds the API key, calls the model, validates the answer against the schema.
Never sees row data or hidden fields.
What leaves your infrastructure
Only schema metadata (field names, types, aliases, enum values) and the user's instruction. Row data never does. The answer is validated on your server and again in the browser before anything is applied.
Install
The server handler and providers are already part of @avinash-baraiya/pragma, so there's nothing extra to install. The Anthropic and AI SDK providers need their SDK (shown below); the others use fetch.
Choose a provider
Every provider takes a model (nothing is hard-coded) and reads your key from wherever you pass it.
// OpenAI, OpenRouter, Groq, Together, Fireworks, Ollama, vLLM, LM Studio, LiteLLM…
import { openAICompatible } from '@avinash-baraiya/pragma/providers/openai-compatible';
const provider = openAICompatible({
baseURL: 'https://api.openai.com/v1', // or http://localhost:11434/v1 for Ollama
model: process.env.OPENAI_MODEL!,
apiKey: process.env.OPENAI_API_KEY,
});// npm install @anthropic-ai/sdk
import { anthropic } from '@avinash-baraiya/pragma/providers/anthropic';
const provider = anthropic({
model: 'claude-haiku-4-5', // a small, fast tier is usually enough
apiKey: process.env.ANTHROPIC_API_KEY,
});import { gemini } from '@avinash-baraiya/pragma/providers/gemini';
const provider = gemini({
model: process.env.GEMINI_MODEL!, // e.g. gemini-3.5-flash-lite
apiKey: process.env.GEMINI_API_KEY!,
});// npm install ai @ai-sdk/mistral (or any AI SDK provider)
import { aiSdk } from '@avinash-baraiya/pragma/providers/ai-sdk';
import { mistral } from '@ai-sdk/mistral';
const provider = aiSdk(mistral('mistral-small-latest'));import { customProvider } from '@avinash-baraiya/pragma';
const provider = customProvider('my-gateway', async (request) => {
const res = await fetch('https://llm.internal/generate', {
method: 'POST',
body: JSON.stringify(request),
signal: request.signal,
});
return { text: await res.text() };
});Pass an array to get a fallback chain: provider: [primary, cheaperFallback].
Mount the server handler
// app/api/pragma/route.ts
import { createPragmaHandler } from '@avinash-baraiya/pragma/server';
import { customersSchema } from '@/lib/schema';
export const POST = createPragmaHandler({
schemas: { customers: customersSchema }, // registered on the server, never sent by the client
provider,
authorize: async ({ request }) => (await getSession(request)) !== null,
});import express from 'express';
import { createPragmaHandler, toNodeHandler } from '@avinash-baraiya/pragma/server';
const app = express();
app.use(express.json());
app.post(
'/api/pragma',
toNodeHandler(createPragmaHandler({ schemas: { customers: customersSchema }, provider })),
);const handler = createPragmaHandler({ schemas: { customers: customersSchema }, provider });
app.post('/api/pragma', (c) => handler(c.req.raw));The handler enforces body and instruction size limits, returns RFC 9457 problem responses for transport errors, and includes a circuit breaker: while the model keeps failing, users still get deterministic answers.
Point the browser at it
import { createEngine } from '@avinash-baraiya/pragma';
import { remoteInterpreter } from '@avinash-baraiya/pragma/providers/remote';
export const engine = createEngine({
schema: customersSchema,
timezone: Intl.DateTimeFormat().resolvedOptions().timeZone,
interpreter: remoteInterpreter({
url: '/api/pragma',
headers: () => ({ authorization: `Bearer ${getToken()}` }),
}),
});Nothing else changes: the same AskBar, chips and table now handle free phrasing too.
Latency and cost
- Answers are cached per instruction, schema, current state and day.
- The system prompt is identical on every request, so providers with prompt caching reuse about 97% of input tokens.
- Pick a small, fast model first and measure it with
pnpm evalin the repository. SettimeoutMs(default 15 s) and a fallback provider so a slow model never blocks users.
Measured results and every option: AI models reference.