Skip to content
AI reference

AI reference

Since v0.3.0

The API of package ai. See Add AI to your app for a walkthrough, AI for the design, and Configuration for the AI_* settings.

Setup

APIDoesSince
ai.ForApp(app, drivers...)*ai.Client from the AI_* settings, with the driver AI_PROVIDER names (fake is built in), in every context the app creates. Each tool call becomes a unit of work (anetos.Unit{Kind: "tool"})v0.3
ai.New(provider, defaults...)A client by hand; defaults (options) apply before each call’s own. Logs to slog.Default()v0.3
ai.WithClient(ctx, c), ai.From(ctx)Put a client in a context; get it (ai.ErrNoClient if absent)v0.3
c.Provider()The client’s ai.Providerv0.3
c.Fake(replies...)Replace the provider with an *ai.Fake (or add replies to it); for testsv0.3
ai.Driver{Name, Open}A provider’s driver: Open(app, cfg) (ai.Provider, error) reads its own settings; a provider that is an io.Closer is closed at shutdownv0.3

Calls

APIReturnsSince
ai.Generate(ctx, prompt, opts...)*ai.Result: the answer after any tool callsv0.3
ai.GenerateObject[T](ctx, prompt, opts...)T, *ai.Result: the answer decoded into T (a struct) and checked with its validate tags; one retry with the problems, then *ai.OutputErrorv0.3
ai.Stream(ctx, prompt, opts...)iter.Seq2[ai.Event, error]: the call’s events; breaking out stops itv0.3
agent.Prompt(ctx, prompt, opts...), agent.Stream(…)Generate and Stream with the agent’s settingsv0.3

The prompt may be empty when ai.Messages gives the conversation.

Options

Applied in order: the client’s defaults, an agent’s, the call’s.

OptionDoesDefault
ai.System(text)Adds system instructions; several are joined by blank lines, in the order the options come (an agent’s Instructions where the agent is)none
ai.Model(name)The model, by the provider’s name for itAI_MODEL
ai.MaxTokens(n)The longest answer, in tokensAI_MAX_TOKENS
ai.Temperature(t)Sampling temperaturethe provider’s
ai.Timeout(d)Each request’s timeout (in a stream, including the reader’s time)AI_TIMEOUT
ai.Messages(msgs...)Adds the conversation before the prompt (a Result’s Messages)none
ai.Tools(tools...)Adds toolsnone
ai.MaxSteps(n)The most requests in a call with tools; then ai.ErrMaxSteps. GenerateObject’s retry is a second call, with its own limitai.DefaultMaxSteps (10)
ai.ProviderOptions(v)The driver’s own request options (a type it defines)none
ai.ForUser(userID)The user the call is for: its usage records and budget (TrackUsage)the signed-in user, if any; a conversation’s user
ai.Using(c)Use client c, not the context’sthe context’s
an ai.AgentIts settings

ai.Agent

FieldTypeDoes
NamestringNames it in logs (agent=)
InstructionsstringIts system instructions
Tools[]ai.ToolIts tools
ModelstringIts model; empty: the default
MaxStepsint0: ai.DefaultMaxSteps
Options[]ai.OptionMore options, before each call’s

Results

TypeFields and methods
ai.ResultMessages (the whole conversation: given messages, prompt, answers, tool results; after an error it can end with tool calls without results, which providers refuse to continue from), Steps ([]*ai.Response), Usage (the total); Text() (the last response’s), Response() (the last)
ai.ResponseMessage, Stop, Usage, Model (that answered), Raw (the provider’s own response); Text(), ToolCalls()
ai.UsageInputTokens (including cached), OutputTokens, CacheReadTokens, CacheWriteTokens; Add(u)
ai.MessageRole (ai.RoleUser, RoleAssistant, RoleTool), Parts (ai.Text, ai.ToolCall{ID, Name, Input}, ai.ToolResult{CallID, Name, Content, IsError}, ai.Reasoning{Provider, Text, Data}: a model’s reasoning its provider needs back, before the part it belongs to; other providers leave it out); Text() (the text parts only), ToolCalls(); JSON {"role", "parts": [{"type": "text" | "tool_call" | "tool_result" | "reasoning", …}]}. ai.UserMessage(text), ai.AssistantMessage(text)
StopMeans
ai.StopEndThe answer is complete
ai.StopToolCallsIt called tools
ai.StopMaxTokensCut off at MaxTokens (not an error for Generate; an OutputError for GenerateObject)
ai.StopRefusalThe model declined (an OutputError for GenerateObject, not retried)
ai.StopOtherA provider-specific reason (see Raw)

Stream events

KindFieldWhen
ai.EventTextTextA piece of the answer
ai.EventToolCallToolCallThe model called a tool
ai.EventToolResultToolResultThe tool’s result
ai.EventResponseResponseA model response is complete (one per step)
ai.EventDoneResultLast, after a successful call

An error is yielded once, and ends the stream.

Tools

APIDoes
ai.Func(name, description, fn)A tool running fn func(ctx, In) (Out, error). In is a struct: its schema is the tool’s input; the model’s input is decoded (*ai.InputError if it doesn’t fit) and validated before fn runs. Out is sent as JSON (a string type as is). Panics on a bad name (1–64 letters, digits, _, -), a non-struct In or bad tags. A panic in fn isn’t recovered
ai.ToolThe interface: Definition() ai.ToolSpec{Name, Description, Input}, Call(ctx, json.RawMessage) (string, error)
A tool’s errorThe model getsThe call
A 4xx status (web.StatusCoder): *ai.InputError, web.Error, auth.ErrForbidden, db.ErrNotFound…IsError result, what a web client would see: the status text; the message and fields of the HTTPError that has the status, or which input field has the wrong JSON type; else the field messages of a validation error in the chainContinues
Any otherNothingStops; returns ai: tool <name>: <error>
A name no tool hasIsError: “There is no tool named …”Continues

Schemas

ai.SchemaFor[T]() builds *ai.Schema (JSON Schema, properties in field order) from struct T, once per type:

GoJSON Schema
structobject of its exported fields by json name, resolved as encoding/json does (json:"-" skips; embedded structs flatten, a shallower or tagged field hides others of its name, ambiguous ones are left out); additionalProperties: false
string, boolstring, boolean
integers, floatsinteger (unsigned: minimum: 0), number; with json:",string", string
slice, arrayarray of the element; []byte: string
map[K]V (string, integer or text keys)object with additionalProperties V
pointerthe element’s type, or null
time.Timestring, date-time; other encoding.TextMarshalers: string (their rules aren’t in the schema)
json.Numbernumber
interface, json.RawMessageany value
a type with its own MarshalJSON or UnmarshalJSONan error: its shape can’t be known
TagAdds
description:"…"description
validate:"required"the field to required, not null, and at least one character, item or key
min, max, size, betweenminLength/maxLength (strings), minimum/maximum (numbers), minItems/maxItems (slices), minProperties/maxProperties (maps)
in:a,benum (of each element, on slices)
email, url, uuid, date, datetime, ipv4, ipv6format (email, uri, uuid, date, date-time, ipv4, ipv6)

Other rules aren’t in the schema, but are checked on the answer. Recursive types are an error.

Errors

ErrorWhenWeb status
ai.ErrNoClientNo client in the context and no ai.Using500
ai.ErrMaxStepsThe model still called tools at the last step500
*ai.BudgetError (UserID, RetryAfter)The call’s user spent their budget; wrapped in a *web.HTTPError whose message tells them when to try again429
ai.ErrConversationChangedAnother call added to the conversation during this one; nothing was stored409
*ai.OutputError (Type, Text, Stop, Err)GenerateObject’s answer was invalid twice, cut off, or refused; Err is a *validate.Errors for broken rules502
ai: <provider>: <error>The provider failed (wraps its error)500; 503 on a timeout

Conversations

APIDoesSince
ai.Migrations()The ai_conversations, ai_messages and ai_usage tables, for migrate.ForAppv0.3
ai.StartConversation(ctx, userID, title)Stores a new *ai.Conversation of the user (an AuthID, up to 100 bytes)v0.3
ai.FindConversation(ctx, userID, id)The user’s conversation, else db.ErrNotFound (404)v0.3
ai.Conversations(ctx, userID)The user’s conversations, the latest changed first, without messagesv0.3
conv.Messages(ctx)Its messages, oldest firstv0.3
conv.Add(ctx, msgs...)Stores messages at the end, without calling a model (the user’s question)v0.3
conv.Prompt(ctx, prompt, opts...), conv.Stream(…)Generate and Stream after its messages; the prompt and answers are stored if the call succeeds (before EventDone)v0.3
conv.Reply(ctx, opts...), conv.StreamReply(…)Answer its last message (the user’s); stored the same wayv0.3
conv.QueueReply(ctx, agent)Answers its last message from a queue job, with the agent registered under agent.Name, as its user (auth.ActAs, with the abilities of the API token ctx was signed in with, if any); Status is ai.StatusQueued, then "", or ai.StatusFailed with Error. A retry runs the tools again; a job that finds the conversation changed does nothing. With the sync queue driver, it runs in the calling request, after its commitv0.3
conv.Delete(ctx)Deletes it and its messages (its usage records stay)v0.3
ai.QueueAgents(app, agents...)Registers the ai.reply job type on the app’s queue, with the agents queued replies may use (by Name); needs queue.ForApp first. Jobs time out after AI_QUEUE_TIMEOUTv0.3

ai.Conversation fields: ID, UserID, Title (cut at 255 bytes), Status, Error, CreatedAt, UpdatedAt. Calls on a conversation are for its user (ForUser), and their usage records name it; ai.Messages options are ignored.

Streaming to the browser

APIDoesSince
ai.SSE(c, events)Writes a stream (ai.Stream, conv.StreamReply…) as server-sent events, through c.Events() (no request or write timeout): text (a piece, HTML-escaped), tool (a tool’s name), error (a 4xx error’s message, else a general one; others are logged), then done, always last; a comment every 15 seconds keeps proxies from closing a quiet streamv0.3

Usage and budgets

APIDoesSince
c.TrackUsage(ai.UsageConfig{Prices, Budget})Records each response in ai_usage and enforces budgets; call at startup. Budgets need the app’s cache. Records are written with the call’s context: inside a transaction that rolls back they go with it, so don’t call models inside transactionsv0.3
ai.Price{Input, Output, CacheRead, CacheWrite}Per million tokens; cache prices of 0 charge Input. p.Cost(usage)v0.3
UsageConfig.Pricesmap[string]ai.Price by model name: the response’s model, else the requested onev0.3
UsageConfig.Budgetfunc(ctx, userID) (ai.Budget, error), once per call with a userv0.3
ai.Budget{Tokens, Cost, Per}Input and output tokens, or cost, per period (aligned to the clock: a day starts at midnight UTC); 0 for no limit. Checked before each request; a response can go past it. Changing a user’s limit starts their count for the period overv0.3
ai.UsageRecordA row of ai_usage: UserID, ConversationID, Agent, Provider, Model, the four token counts, Cost, Estimated, CreatedAt. A stream the reader left partway, or a request cut off by a timeout or a cancellation, is recorded and counted too, estimated at about four bytes a token for its input and the output it yielded (Estimated)v0.3
ai.TotalUsage(ctx, userID, since)ai.UsageTotal{Usage, Cost, Responses} since a timev0.3

Embeddings

See Search by meaning.

APIDoesSince
ai.Embed(ctx, dims, texts...)[]ai.Vector: the texts’ embeddings as documents, with the client’s embedding model (AI_EMBEDDING_MODEL), in requests of at most 96 texts; dims asks for vectors of that size (0: the model’s), and vectors of another size are an error. Each request’s usage is recorded (agent embed) and counts against budgetsv0.3
ai.EmbedQuery(ctx, dims, text)The embedding of a query (what to find documents with)v0.3
ai.Vectordb.Vector, a []float32v0.3
ai.EmbedderEmbed(ctx, *ai.EmbedRequest) (*ai.EmbedResponse, error): the OpenAI (and compatible), Gemini and fake providers have it; Anthropic’s doesn’t. EmbedRequest{Model, Inputs, Dimensions, Purpose} (ai.EmbedForDocument, ai.EmbedForQuery); EmbedResponse{Vectors, Usage, Model, Estimated}v0.3
ai.ErrNoEmbedderThe client has no embeddings providerv0.3
client.EmbeddingModel(), client.SetEmbedder(e, model)The embedding model’s name; set the embeddings provider of a client made with ai.Newv0.3
ai.EmbeddingsFor(app, ai.EmbeddingsConfig[T]{Text, Dimensions, FixedSize, ChunkSize, Title, Scope})*ai.Embeddings[T], which keeps the embeddings of T’s records in <table>_embeddings (migrate.Schema.CreateEmbeddings). Text is a record’s text, Dimensions the vectors’ size (the table’s, which the model is asked for; with FixedSize, for models that make one size, only checked), ChunkSize the most characters in a chunk (2000), Title names records in the tool’s results, Scope narrows every search (with the search’s context). After ai.ForApp; with the app’s queue (queue.ForApp first), registers the job ai.embed:<table>. Adds the command ai:embed [table…]v0.3
e.Sync(ctx, rows...)Updates the rows’ embeddings: in queue jobs of a hundred rows dispatched after the transaction commits, or right away without a queue (after the commit, inside a transaction)v0.3
e.SyncNow(ctx, rows...)Updates them now: splits the text into chunks (paragraphs, sentences, words), embeds those whose text or model changed, and replaces the record’s chunks. Not inside a transactionv0.3
e.SyncAll(ctx)Deletes the chunks of deleted records (db.PruneChunks), then syncs every record, a hundred at a time (what ai:embed runs)v0.3
e.Search(ctx, query, limit, scopes...)[]ai.Passage[T]{Record, Text, Position, Distance}: the records nearest the query, best first (at most limit; 10 for 0), with their nearest chunk; hybrid (db.Q.Hybrid) when the table has a search index, else db.Q.Similar; a blank query finds nothing. A record found only by its words has Position -1, the start of its text, and Distance 1v0.3
e.Tool(name, description, limit)An ai.Tool that searches: the model sends {"query": …} and gets []ai.SearchResult{ID, Title, Text}v0.3

Providers

PackageAI_PROVIDERConstructorOptions (for ai.ProviderOptions)Response.Raw
drivers/anthropicanthropicanthropic.Driver(), anthropic.New(key, opts...)ThinkingBudget (extended thinking), Params func(*anthropic.MessageNewParams)*anthropic.Message
drivers/openaiopenaiopenai.Driver(), openai.New(key, opts...)ReasoningEffort, Params func(*openai.ChatCompletionNewParams)*openai.ChatCompletion (streams: []openai.ChatCompletionChunk)
drivers/openaiopenai-compatibleopenai.CompatibleDriver(), openai.NewCompatible(url, key, opts...)the samethe same
drivers/geminigeminigemini.Driver(), gemini.New(ctx, genai.ClientConfig{APIKey: key, …})ThinkingBudget *int32, Config func(*genai.GenerateContentConfig)*genai.GenerateContentResponse (streams: a slice of them)

Each provider’s Client() returns its SDK client. All use their API’s chat endpoint (OpenAI’s: Chat Completions, which compatible servers have) and the provider’s structured output; the compatible provider sends max_tokens, OpenAI’s max_completion_tokens. Rate limits and server errors are retried twice (by the Anthropic and OpenAI SDKs; the Gemini driver turns the SDK’s retries on). The Anthropic and compatible providers send nothing from the SDKs’ environment variables; OpenAI’s reads OPENAI_ORG_ID, OPENAI_PROJECT_ID and the like, which are yours.

ProviderJSON Schema it takes for structured outputOther rules
Anthropic (no maps or any fields: refused before sending; at most 24 optional fields and 16 nullable or union ones)types, enum, format; nullable as anyOfIn the description
OpenAI (strict mode, when no maps or any fields)types, enum; every property required, optional ones nullableIn the description
Geminitypes, enum, format (date-time, date, time), minimum, maximum, minItems, maxItems; nullable as anyOfIn the description

Tool inputs get all of JSON Schema, except on Gemini (its dialect, as above).

Embeddings: OpenAI’s and compatible servers’ Embeddings API (dimensions when asked), and Gemini’s (RETRIEVAL_DOCUMENT or RETRIEVAL_QUERY, outputDimensionality; Gemini doesn’t count tokens, so its usage is estimated at four bytes a token). Anthropic has no embeddings API: set AI_EMBEDDING_PROVIDER.

Writing a provider

APIDoes
ai.ProviderName(), Generate(ctx, *ai.Request) (*ai.Response, error), Stream(ctx, *ai.Request) iter.Seq2[ai.Event, error] (EventTexts and EventToolCalls, then one EventResponse with the whole response)
schema.Map(ai.SchemaOptions{Keywords, Formats, AllRequired, NullableAnyOf})The schema as a map[string]any in the provider’s dialect: the constraint keywords it takes (of ai.ConstraintKeywords) and the string formats (all, if Formats is empty), the others in words in the description; every property required (optional ones nullable); nullable as anyOf. Properties keep their order when marshaled (a provider may still reorder them: Anthropic puts required ones first)
aitest.Run(t, aitest.Config{Name, Model, KeyEnv, New, Dir, Skip, EmbeddingModel})The conformance suite (package ai/aitest): text, streams, a conversation, tool calls and their results (streamed too), structured output, the token limit, errors, and, with EmbeddingModel, embeddings (Embed). It replays recorded HTTP exchanges (testdata/aitest/<Test>.json) and checks the requests match; ANETOS_AI_RECORD=1 records them from the provider (its key in KeyEnv; headers are never saved), ANETOS_AI_LIVE=1 calls it without recording, ANETOS_AI_UPDATE_REQUESTS=1 rewrites the recorded requests. ANETOS_TEST_<NAME>_MODEL (and _EMBEDDING_MODEL) picks the model to record with
ai.RequestModel (the drivers refuse an empty one), System, Messages, Tools ([]ai.ToolSpec), Output (*ai.OutputSpec{Name, Schema}: structured output), MaxTokens, Temperature (*float64), Options (the driver’s own type, or ignore); Prompt()
ai.Fake, ai.NewFake(replies...)The scripted provider, safe for concurrent use: Add, Requests, Remaining, Embeddings (its embedding requests: it embeds texts by their words, without replies); replies ai.FakeText, FakeObject, FakeToolCall, FakeError, or an ai.FakeReply function
ai.FakeDriver()AI_PROVIDER=fake