Sessions
Create a session against a real model, then steer, queue, withdraw and compact while it runs.
A real model session
Run from the repository root. Set RNA_MODEL_BASE_URL, RNA_MODEL_ID and RNA_MODEL_API_KEY; RNA_MODEL_PROTOCOL is openai (base URL includes /v1) or anthropic. This makes real model requests.
import { resolve } from 'node:path';
import { createSession, createWorkspaceTools } from './packages/sdk/src/index.mjs';
const model = {
providerId: 'configured',
protocol: process.env.RNA_MODEL_PROTOCOL || 'openai',
baseUrl: process.env.RNA_MODEL_BASE_URL,
modelId: process.env.RNA_MODEL_ID,
contextWindow: Number(process.env.RNA_CONTEXT_WINDOW || 32000),
maxOutputTokens: Number(process.env.RNA_MAX_OUTPUT_TOKENS || 2048),
reasoning: 'off',
cacheRetention: 'short',
};
const cwd = process.cwd();
const stateDir = resolve('.rna-sdk-state');
const session = await createSession({
sessionId: 'example-conversation',
projectId: 'example-project',
cwd,
stateDir,
model,
apiKey: process.env.RNA_MODEL_API_KEY,
tools: await createWorkspaceTools(cwd, { deniedPaths: [stateDir] }),
systemPrompt: 'You are Rna Agent. Help with the readable project content; never claim actions you did not take.',
onEvent(event) {
if (event.type === 'message_update' && event.assistantMessageEvent?.type === 'text_delta') {
process.stdout.write(event.assistantMessageEvent.delta);
}
},
});
try {
await session.prompt('Look at the top-level folders and say what deserves attention next.');
console.log('\n', session.snapshot().usage);
} finally {
await session.dispose();
}The sizes are this example’s request budget, not provider specs. Creating a session again with the same stateDir + sessionId + projectId reads the same JSONL log.
Input while running
const running = session.prompt('Start the check');
const queued = await session.followUp('Suggest tests when the check is done');
await session.steer('Compatibility first, everything else after');
await session.withdraw(queued.id); // possible while still queued
await running;steerenters the context at a model or tool safe boundary. It cannot change an HTTP request already sent or undo executed side effects.followUpcontinues once the current work settles.abort()requests a stop;resume()continues the durable session.yieldAtBoundary({ requestId })saves a cooperative pause, applied after the current request or tool batch settles.
Compaction
Automatic compaction is on by default: near 80% of the configured context, after a complete tool batch, older history is summarized in sections while the last two reply and tool groups are kept. If the summary fails or does not shrink the context, the original is kept and the turn stops.
Compaction follows pi, so a conversation can continue on a model with a smaller window: the summary reads a plain transcript where tool results, call arguments, host context and reasoning keep their first 2000 characters; a batch the service reports as too long is retried at half size (up to three times); an answer refused as too long is compacted to half of what was sent and retried once. Overflow is recognized from provider error text, Codex detail bodies and response.failed stream events.
Manual compaction keeps the raw log and needs a real summary from you:
await session.compact({
keepLastTurns: 2,
summarize: async ({ messages, previousSummary, instructions }) => {
return await yourSummarizer({ messages, previousSummary, instructions });
},
});yourSummarizer is implemented by the host; it is not exported by the package.
Hooks at request boundaries
| Hook | When | Use |
|---|---|---|
beforeRequest | After the previous tool batch, before the next request | Return a full new tool set or sourced context updates |
beforeCompletion | Before every real model request, including compaction | Budget reservation |
beforeFinish / afterRun | Around settling | Record results; never a new permission source |
beforeTool / afterTool | After a tool call's arguments are validated and before it runs, and after it ran | The host's tool policy: can refuse the call or attach a note to the result; never a new permission source |
Limits and retries
streamCompletion defaults to a 120 s idle limit and 15 min total per request; SSE heartbeats only refresh the idle timer. Automatic retries happen only on explicit HTTP 429 or 5xx before any SSE response was accepted, at most twice, honoring Retry-After up to 60 s. Connection errors, half streams and executed tools are never replayed implicitly.
Anthropic prompt cache
On the Anthropic protocol, history is append-only (appendOnlyActive) and thinking is bound to the conversation. cacheKeepAlive(input) re-sends the previous request once with max_tokens: 0 and no streaming: it prefills only and produces no output, yet refreshes the cache timer. Requests with thinking.type: "enabled" or structured output cannot be kept warm this way and are refused. It never retries; the host decides when it is worth calling.
Images
Pass images as { type: 'image', data, mimeType } through session.prompt(text, { images }). The model must declare input: ['text', 'image']; otherwise you get a clear diagnostic, never a silent model switch or a pretend read.