Compare commits
3 Commits
5166fde764
...
939598a187
| Author | SHA1 | Date | |
|---|---|---|---|
| 939598a187 | |||
| db4803c557 | |||
| 797ee5f2d2 |
176
docs/deploy-db4803c.md
Normal file
176
docs/deploy-db4803c.md
Normal file
@@ -0,0 +1,176 @@
|
||||
# Deploying `db4803c` and switching the model vendor to Gemini
|
||||
|
||||
Two things land together, and the order matters: the image **must** be
|
||||
running before the configuration switches vendor. Old binary on Gemini
|
||||
config = every tool-using run dies on its second model call (§2). New binary
|
||||
on Groq config = works exactly as today. So: image first, config second.
|
||||
|
||||
Live today: Groq free tier, `8000 TPM`, **35% of runs since Sep 9 end in
|
||||
`gateway.rate_limited`**. That number is why this deploy exists.
|
||||
|
||||
---
|
||||
|
||||
## 1. What is being deployed
|
||||
|
||||
`db4803c`, on `origin/main`. Five commits since `8e36faf`:
|
||||
|
||||
| Commit | What |
|
||||
| --- | --- |
|
||||
| `822b3b1` | `.env` untracked again; ignore rules `3455ad0` deleted are back |
|
||||
| `b765495` | Archiving an agent a published agent delegates to → 409 |
|
||||
| `5166fde` | `GatewayFailure` termination; **migration 000016** |
|
||||
| `797ee5f` | Provider metadata round-tripped on tool calls — **Gemini needs this** |
|
||||
| `db4803c` | An unsaved trajectory is logged, not just noted in itself |
|
||||
|
||||
**One migration.** `000016` widens `agent_runs.termination_check` to admit
|
||||
`GatewayFailure`. The `migrate` init container applies it before the API
|
||||
starts. Its down migration is verified (folds rows to `ToolFailure` before
|
||||
narrowing the CHECK), so a rollback of the image is safe.
|
||||
|
||||
## 2. Why the image must go first
|
||||
|
||||
Gemini 3 models attach a `thought_signature` to every function call and
|
||||
reject the follow-up request without it. `797ee5f` teaches the gateway to
|
||||
carry it back. The image on the cluster today does not, and it was proved on
|
||||
2026-09-22: pointed at Gemini, the first model call succeeded, the tool ran,
|
||||
the second call answered `400 Function call is missing a thought_signature`.
|
||||
Rolled back to Groq within minutes.
|
||||
|
||||
## 3. What was verified before writing this
|
||||
|
||||
With the `797ee5f` binary, locally, against the production database and the
|
||||
real Gemini API through a request-logging proxy:
|
||||
|
||||
| Check | Result |
|
||||
| --- | --- |
|
||||
| `positions-agent`, 2 model calls | `Completed`, signature present on the echoed call |
|
||||
| `krow-workforce-agent`, 3 tool-calling turns, **12,123 tokens** | `Completed`, correct answer. This exceeds Groq's entire per-minute ceiling |
|
||||
| Streaming path carries the signature | unit test + live run |
|
||||
| `GatewayFailure` reaches the surface with its own wording | seen live before rollback |
|
||||
| Full Go suite against a real Postgres (the DB tests skip without one) | 18/18 packages |
|
||||
| Migration 000016 up → down → up on a scratch database | clean |
|
||||
|
||||
Model reliability, six bare calls each, 2026-09-22 ~13:00 IST:
|
||||
|
||||
| Model | HTTP codes |
|
||||
| --- | --- |
|
||||
| `gemini-3.8-flash` | 503 503 200 200 503 503 |
|
||||
| `gemini-3.5-flash` | 503 200 503 200 200 200 |
|
||||
| `gemini-3.5-flash-lite` | 200 200 200 200 200 200 |
|
||||
|
||||
The gateway retries a 503 three times with backoff; at the rates above a
|
||||
three-call run on either larger model still fails often. **All three tiers
|
||||
run `gemini-3.5-flash-lite`** until the larger models stop shedding load or
|
||||
the key is on a paid tier. `gemini-3.1-pro-preview` answers 429 (pro is not
|
||||
on the free tier); the 2.5 family is listed but blocked for new keys.
|
||||
|
||||
## 4. Build and push the image (the other machine)
|
||||
|
||||
The Dockerfile cross-compiles, so any host with `buildx` and a Docker Hub
|
||||
login works:
|
||||
|
||||
```bash
|
||||
git checkout db4803c
|
||||
docker buildx build --platform linux/amd64 \
|
||||
-f infrastructure/Dockerfile.api \
|
||||
-t doormile/krowbackend:db4803c -t doormile/krowbackend:latest \
|
||||
--push .
|
||||
```
|
||||
|
||||
Two tags on purpose: `:latest` is what the StatefulSet pulls; `:db4803c` is
|
||||
what you roll back **to** if you need to (§7). Confirm before touching the
|
||||
cluster:
|
||||
|
||||
```bash
|
||||
docker buildx imagetools inspect doormile/krowbackend:latest | grep -E 'Platform|Digest' | head -3
|
||||
```
|
||||
|
||||
## 5. Switch the cluster (on the server, as root)
|
||||
|
||||
The Gemini key is already staged in the Secret as `MODEL_API_KEY_GEMINI`;
|
||||
the Groq key stays as `MODEL_API_KEY_GROQ`. The manifests in
|
||||
`/opt/kubernetes/manifests/krow` already describe the Gemini configuration
|
||||
(committed `pending`), so the config half is an `apply`.
|
||||
|
||||
```bash
|
||||
# 1. the credential the API reads becomes the Gemini one
|
||||
kubectl -n krow patch secret krow-model --type=json \
|
||||
-p '[{"op":"copy","from":"/data/MODEL_API_KEY_GEMINI","path":"/data/MODEL_API_KEY"}]'
|
||||
|
||||
# 2. configmap → Gemini base URL and model ids
|
||||
kubectl apply -k /opt/kubernetes/manifests/krow/
|
||||
|
||||
# 3. new pods: pull :latest, run migration 16, boot on the new config
|
||||
kubectl -n krow rollout restart statefulset/krow
|
||||
kubectl -n krow rollout status statefulset/krow --timeout=5m
|
||||
```
|
||||
|
||||
`rollout status` waits for krow-2, then krow-1, each gated on readiness. If
|
||||
krow-2 does not come up, krow-1 is still serving on the old image and Groq.
|
||||
|
||||
## 6. Smoke test
|
||||
|
||||
```bash
|
||||
# migration 16 applied?
|
||||
kubectl -n krow logs krow-2 -c migrate | tail -2 # want: 16/u gateway_failure_termination
|
||||
|
||||
# boot on the right vendor?
|
||||
kubectl -n krow logs krow-2 -c api | grep -m1 '"listening"' | grep -o '"endpoints":[0-9]*' # want 71
|
||||
|
||||
# a real run, through the public URL
|
||||
T=$(curl -s -D - -o /dev/null -X POST https://mcp.krowforce.com/api/v1/auth/login \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"email":"demo@krow.app","password":"<demo password>"}' \
|
||||
| sed -n 's/^[Ss]et-[Cc]ookie: krow_session=\([^;]*\).*/\1/p')
|
||||
curl -s -X POST https://mcp.krowforce.com/api/v1/agents/65bfd77d-2f74-4548-ab52-4e720e153397/runs \
|
||||
-H "Cookie: krow_session=$T" -H 'Content-Type: application/json' \
|
||||
-d '{"input":"How many open positions are there?"}' | grep -E '"(termination|output)"'
|
||||
```
|
||||
|
||||
Want `"termination": "Completed"` and a count. `GatewayFailure` with
|
||||
"usually it is busy" is Gemini shedding load — retry once. `ToolFailure` on
|
||||
the **second** model call means the old image is still running (§2).
|
||||
|
||||
Then check the trajectory landed with the right model:
|
||||
|
||||
```bash
|
||||
# from a machine with psql / the postgres image; DATABASE_URL from secret/krow-db
|
||||
psql "$DATABASE_URL" -Atc "SELECT model, termination FROM agent_runs ORDER BY started_at DESC LIMIT 1"
|
||||
```
|
||||
|
||||
Want `gemini-3.5-flash-lite | Completed`.
|
||||
|
||||
## 7. Rollback
|
||||
|
||||
Config only (image stays — it works on Groq too):
|
||||
|
||||
```bash
|
||||
kubectl -n krow patch secret krow-model --type=json \
|
||||
-p '[{"op":"copy","from":"/data/MODEL_API_KEY_GROQ","path":"/data/MODEL_API_KEY"}]'
|
||||
kubectl -n krow patch cm krow-config --type merge -p '{"data":{
|
||||
"MODEL_BASE_URL":"https://api.groq.com/openai/v1",
|
||||
"MODEL_FAST":"openai/gpt-oss-20b","MODEL_BALANCED":"openai/gpt-oss-120b","MODEL_DEEP":"openai/gpt-oss-120b"}}'
|
||||
kubectl -n krow rollout restart statefulset/krow
|
||||
```
|
||||
|
||||
Image too (only if `db4803c` itself misbehaves):
|
||||
|
||||
```bash
|
||||
kubectl -n krow set image statefulset/krow api=doormile/krowbackend:<previous tag>
|
||||
```
|
||||
|
||||
Migration 16 stays applied; the old binary never writes `GatewayFailure`,
|
||||
so the wider CHECK is harmless to it. Reverse it only if you must:
|
||||
`migrate ... down 1` — it folds existing `GatewayFailure` rows to `ToolFailure`.
|
||||
|
||||
## 8. Not covered here
|
||||
|
||||
- **`activity-agent` is archived while `krow-workforce-agent v2` delegates to
|
||||
it.** `b765495` prevents this happening again; it does not repair the
|
||||
existing case. Either unarchive `activity-agent` or publish workforce v3
|
||||
without it — a product decision.
|
||||
- **`OAUTH_LOGIN_PATH` (`/login`) 404s on `mcp.krowforce.com`.** A signed-out
|
||||
MCP consent redirect goes nowhere. Signed-in users are unaffected.
|
||||
- **Rotation.** The Anthropic key in `3455ad0`'s history, the Groq key, the
|
||||
Gemini key (pasted in a chat), the DB admin password (8 chars, public IP,
|
||||
no TLS), the root SSH password, the demo login.
|
||||
@@ -85,6 +85,25 @@ type ToolCall struct {
|
||||
ID string
|
||||
Name string
|
||||
Input json.RawMessage
|
||||
|
||||
// Extra is provider metadata attached to the call, carried back to the
|
||||
// provider verbatim on the next turn and never read here.
|
||||
//
|
||||
// It exists because at least one provider requires it. Gemini 3 models
|
||||
// attach a "thought signature" to every function call and REJECT the
|
||||
// follow-up request — 400, "Function call is missing a thought_signature"
|
||||
// — if the assistant message that echoes the call does not carry it back.
|
||||
// A gateway that rebuilds the assistant turn from ID, Name and Input alone
|
||||
// drops it, and every tool-using run dies on its second model call while
|
||||
// the first one looked perfectly healthy. That is exactly what happened
|
||||
// on 2026-09-22 when production was pointed at Gemini.
|
||||
//
|
||||
// The gateway does not know what is in it and must not: the whole point
|
||||
// of speaking one wire shape is that a vendor's private fields pass
|
||||
// through untouched. It is the raw JSON of the call's extra_content
|
||||
// object, or nil when the provider sent none, in which case it is omitted
|
||||
// from the request again.
|
||||
Extra json.RawMessage
|
||||
}
|
||||
|
||||
// ToolResult is what came back, on its way to the model.
|
||||
|
||||
@@ -247,6 +247,7 @@ func (g *OpenAIGateway) decode(
|
||||
// The raw JSON, not a parsed value — handed to the handler's own
|
||||
// decoder rather than matched on as a string here.
|
||||
Input: json.RawMessage(args),
|
||||
Extra: c.ExtraContent,
|
||||
})
|
||||
}
|
||||
|
||||
@@ -301,6 +302,10 @@ type oaiToolCall struct {
|
||||
ID string `json:"id,omitempty"`
|
||||
Type string `json:"type,omitempty"`
|
||||
Function oaiFunctionRef `json:"function"`
|
||||
|
||||
// ExtraContent is the provider's own metadata on the call, round-tripped
|
||||
// as raw JSON. See ToolCall.Extra for why it is not optional.
|
||||
ExtraContent json.RawMessage `json:"extra_content,omitempty"`
|
||||
}
|
||||
|
||||
type oaiFunctionRef struct {
|
||||
@@ -496,6 +501,7 @@ func encodeOpenAIMessages(system string, msgs []Message) []oaiMessage {
|
||||
ID: c.ID,
|
||||
Type: "function",
|
||||
Function: oaiFunctionRef{Name: c.Name, Arguments: args},
|
||||
ExtraContent: c.Extra,
|
||||
})
|
||||
}
|
||||
out = append(out, msg)
|
||||
@@ -549,6 +555,20 @@ func translateOpenAI(status int, body []byte) error {
|
||||
// worth reading and not often enough to depend on, so an unparseable body
|
||||
// yields nothing rather than failing a failure.
|
||||
func openAIErrorMessage(body []byte) string {
|
||||
// Gemini wraps its error in a one-element ARRAY — `[{"error":{...}}]` —
|
||||
// where OpenAI, Groq and the rest send the object bare. Unwrapped here
|
||||
// rather than tolerated as "no detail", because the detail is the whole
|
||||
// value of the field: for two weeks the trajectory said only "the model
|
||||
// rejected the request" when the body said "Function call is missing a
|
||||
// thought_signature", and the difference was a day of diagnosis.
|
||||
body = bytes.TrimSpace(body)
|
||||
if bytes.HasPrefix(body, []byte("[")) {
|
||||
var many []json.RawMessage
|
||||
if err := json.Unmarshal(body, &many); err != nil || len(many) == 0 {
|
||||
return ""
|
||||
}
|
||||
body = many[0]
|
||||
}
|
||||
var envelope struct {
|
||||
Error struct {
|
||||
Message string `json:"message"`
|
||||
@@ -744,6 +764,11 @@ func (a *streamAccumulator) addToolCallDeltas(deltas []oaiToolCall) {
|
||||
if d.Function.Name != "" {
|
||||
call.Function.Name = d.Function.Name
|
||||
}
|
||||
// Provider metadata arrives whole on one fragment, like the id. Kept
|
||||
// when non-empty so a later empty fragment does not erase it.
|
||||
if len(d.ExtraContent) > 0 {
|
||||
call.ExtraContent = d.ExtraContent
|
||||
}
|
||||
// Arguments are the fragmented field: concatenated, never replaced.
|
||||
call.Function.Arguments += d.Function.Arguments
|
||||
}
|
||||
|
||||
@@ -209,3 +209,29 @@ func TestStreamCompleteUsesTheStreamingPath(t *testing.T) {
|
||||
t.Errorf("Text = %q", resp.Text)
|
||||
}
|
||||
}
|
||||
|
||||
// The streamed shape of TestToolCallProviderMetadataIsRoundTripped: the
|
||||
// metadata arrives on one fragment, and later fragments that carry only
|
||||
// argument text must not erase it.
|
||||
func TestStreamKeepsToolCallProviderMetadata(t *testing.T) {
|
||||
const sig = `{"google":{"thought_signature":"El4KXAFpFH0T4CM3"}}`
|
||||
acc, err := accumulateSSE(strings.NewReader(strings.Join([]string{
|
||||
`data: {"choices":[{"delta":{"tool_calls":[{"index":0,"id":"call_a","function":{"name":"open_positions","arguments":""},"extra_content":` + sig + `}]}}]}`,
|
||||
`data: {"choices":[{"delta":{"tool_calls":[{"index":0,"function":{"arguments":"{}"}}]}}]}`,
|
||||
`data: {"choices":[{"delta":{},"finish_reason":"tool_calls"}]}`,
|
||||
`data: [DONE]`,
|
||||
}, "\n\n")), func(string) {})
|
||||
if err != nil {
|
||||
t.Fatalf("accumulateSSE: %v", err)
|
||||
}
|
||||
msg := acc.message()
|
||||
if len(msg.ToolCalls) != 1 {
|
||||
t.Fatalf("got %d tool calls, want 1", len(msg.ToolCalls))
|
||||
}
|
||||
if string(msg.ToolCalls[0].ExtraContent) != sig {
|
||||
t.Errorf("extra_content after streaming = %s, want %s", msg.ToolCalls[0].ExtraContent, sig)
|
||||
}
|
||||
if msg.ToolCalls[0].Function.Arguments != "{}" {
|
||||
t.Errorf("arguments = %q, want {}", msg.ToolCalls[0].Function.Arguments)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -339,3 +339,96 @@ func TestBaseURLDefaultsAndTrimsSlash(t *testing.T) {
|
||||
t.Errorf("endpoint = %q, want the trailing slash collapsed", got)
|
||||
}
|
||||
}
|
||||
|
||||
// THE FAILURE THIS EXISTS FOR: a provider that attaches private metadata to a
|
||||
// tool call and refuses the follow-up without it. Gemini 3 does exactly this
|
||||
// ("Function call is missing a thought_signature"), and a gateway that rebuilt
|
||||
// the assistant turn from id, name and arguments alone killed every tool-using
|
||||
// run on its second model call — after a first call that looked healthy.
|
||||
//
|
||||
// The round trip is tested end to end: the provider's extra_content on the
|
||||
// response must reappear, byte for byte, on the next request's echo of that
|
||||
// call. The gateway must not care what is inside it.
|
||||
func TestToolCallProviderMetadataIsRoundTripped(t *testing.T) {
|
||||
const sig = `{"google":{"thought_signature":"El4KXAFpFH0T4CM3"}}`
|
||||
|
||||
gw, captured := serve(t, func(w http.ResponseWriter, _ *oaiRequest) {
|
||||
_, _ = io.WriteString(w, `{
|
||||
"choices":[{"message":{"role":"assistant","tool_calls":[
|
||||
{"id":"call_x","type":"function",
|
||||
"function":{"name":"open_positions","arguments":"{}"},
|
||||
"extra_content":`+sig+`}]},
|
||||
"finish_reason":"tool_calls"}],
|
||||
"usage":{"prompt_tokens":10,"completion_tokens":5}
|
||||
}`)
|
||||
})
|
||||
|
||||
resp, err := gw.Complete(context.Background(), ask("how many open positions?"))
|
||||
if err != nil {
|
||||
t.Fatalf("Complete: %v", err)
|
||||
}
|
||||
if len(resp.ToolCalls) != 1 {
|
||||
t.Fatalf("got %d tool calls, want 1", len(resp.ToolCalls))
|
||||
}
|
||||
if string(resp.ToolCalls[0].Extra) != sig {
|
||||
t.Fatalf("Extra = %s, want the provider's extra_content verbatim", resp.ToolCalls[0].Extra)
|
||||
}
|
||||
|
||||
// Second turn: the loop echoes the assistant's call and adds the result.
|
||||
// This is the request Gemini rejects when the signature is missing.
|
||||
_, err = gw.Complete(context.Background(), Request{Tier: TierBalanced, Messages: []Message{
|
||||
{Role: RoleUser, Text: "how many open positions?"},
|
||||
{Role: RoleAssistant, ToolCalls: resp.ToolCalls},
|
||||
{Role: RoleUser, ToolResults: []ToolResult{{CallID: "call_x", Content: `{"count":14}`}}},
|
||||
}})
|
||||
if err != nil {
|
||||
t.Fatalf("second Complete: %v", err)
|
||||
}
|
||||
|
||||
var echoed *oaiToolCall
|
||||
for i := range captured.Messages {
|
||||
if len(captured.Messages[i].ToolCalls) > 0 {
|
||||
echoed = &captured.Messages[i].ToolCalls[0]
|
||||
}
|
||||
}
|
||||
if echoed == nil {
|
||||
t.Fatalf("the second request did not echo the assistant's tool call: %+v", captured.Messages)
|
||||
}
|
||||
if string(echoed.ExtraContent) != sig {
|
||||
t.Errorf("echoed extra_content = %s, want %s", echoed.ExtraContent, sig)
|
||||
}
|
||||
}
|
||||
|
||||
// A provider that sends no metadata must not receive an "extra_content": null
|
||||
// it never asked for. Absent stays absent.
|
||||
func TestToolCallWithoutProviderMetadataOmitsTheField(t *testing.T) {
|
||||
msgs := []Message{
|
||||
{Role: RoleAssistant, ToolCalls: []ToolCall{{ID: "call_1", Name: "open_positions", Input: json.RawMessage(`{}`)}}},
|
||||
}
|
||||
raw, err := json.Marshal(encodeOpenAIMessages("", msgs))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if strings.Contains(string(raw), "extra_content") {
|
||||
t.Errorf("extra_content was emitted for a call that had none: %s", raw)
|
||||
}
|
||||
}
|
||||
|
||||
// Gemini wraps its error in a one-element array. The detail must survive,
|
||||
// because a bare "the model rejected the request" is the difference between a
|
||||
// one-line diagnosis and a day of one.
|
||||
func TestProviderErrorDetailSurvivesArrayEnvelope(t *testing.T) {
|
||||
cases := map[string]string{
|
||||
`{"error":{"message":"bare object"}}`: "bare object",
|
||||
`[{"error":{"message":"array wrapped"}}]`: "array wrapped",
|
||||
` [ {"error":{"message":"padded"}} ] `: "padded",
|
||||
`{"message":"top level"}`: "top level",
|
||||
`[]`: "",
|
||||
`not json`: "",
|
||||
}
|
||||
for body, want := range cases {
|
||||
if got := openAIErrorMessage([]byte(body)); got != want {
|
||||
t.Errorf("openAIErrorMessage(%s) = %q, want %q", body, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -162,10 +162,24 @@ func (s *Server) handleAgentRun(w http.ResponseWriter, r *http.Request) {
|
||||
writeError(w, s.log, runLoadError(runErr))
|
||||
return
|
||||
}
|
||||
s.logUnsaved(ident, res)
|
||||
|
||||
writeJSON(w, http.StatusOK, buildRunResponse(res))
|
||||
}
|
||||
|
||||
// logUnsaved is the operator's record of a run whose trajectory did not
|
||||
// persist. The run itself already answered; §6 says the trajectory is not
|
||||
// optional telemetry, so losing one is an error even when nothing else went
|
||||
// wrong, and it carries every field §10 asks a log line to carry.
|
||||
func (s *Server) logUnsaved(ident authctx.Identity, res *runtime.ExecutionResult) {
|
||||
for _, detail := range res.Unsaved {
|
||||
s.log.Error("trajectory unsaved",
|
||||
"run_id", res.RunID, "tenant_id", ident.OrgID,
|
||||
"agent_key", res.AgentID, "agent_version", res.AgentVersion,
|
||||
"detail", detail)
|
||||
}
|
||||
}
|
||||
|
||||
// buildRunResponse turns a runtime result into the client's shape.
|
||||
//
|
||||
// Every termination answers 200. That looks wrong at first and is not: the
|
||||
@@ -336,6 +350,7 @@ func (s *Server) streamAgentRun(w http.ResponseWriter, r *http.Request, ident au
|
||||
writeError(w, s.log, runLoadError(runErr))
|
||||
return
|
||||
}
|
||||
s.logUnsaved(ident, res)
|
||||
writeJSON(w, http.StatusOK, buildRunResponse(res))
|
||||
return
|
||||
}
|
||||
@@ -383,6 +398,7 @@ func (s *Server) streamAgentRun(w http.ResponseWriter, r *http.Request, ident au
|
||||
return
|
||||
}
|
||||
|
||||
s.logUnsaved(ident, res)
|
||||
send(map[string]any{"run": buildRunResponse(res)})
|
||||
fmt.Fprint(w, "data: [DONE]\n\n")
|
||||
flusher.Flush()
|
||||
|
||||
@@ -605,8 +605,10 @@ func (m *ModelExecutor) finish(
|
||||
// A sink that fails must not fail the run — the answer was already
|
||||
// produced. It is recorded in the trajectory we could not save, which is
|
||||
// the best available place for it.
|
||||
var unsaved []string
|
||||
if err := m.sink.Save(ctx, traj); err != nil {
|
||||
rec.Error("runtime.trajectory_unsaved", err.Error())
|
||||
unsaved = append(unsaved, traj.RunID+": "+err.Error())
|
||||
}
|
||||
|
||||
// Delegated runs are written AFTER this one, because parent_run_id is a
|
||||
@@ -617,10 +619,12 @@ func (m *ModelExecutor) finish(
|
||||
if err := m.sink.Save(ctx, child); err != nil {
|
||||
rec.Error("runtime.subrun_unsaved",
|
||||
fmt.Sprintf("%s: %s", child.RunID, err.Error()))
|
||||
unsaved = append(unsaved, child.RunID+": "+err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
res := &ExecutionResult{
|
||||
Unsaved: unsaved,
|
||||
Success: term == TerminationCompleted,
|
||||
Output: output,
|
||||
AgentID: agent.ID,
|
||||
|
||||
@@ -265,6 +265,28 @@ func TestSinkFailureDoesNotFailTheRun(t *testing.T) {
|
||||
if res.Output != "the answer" {
|
||||
t.Errorf("Output = %q, want the answer through", res.Output)
|
||||
}
|
||||
// ...but the loss must be reported somewhere that survives. The runtime
|
||||
// has no logger; the surface reads this and writes the operator's line.
|
||||
// Before this field existed the only record was an entry in the very
|
||||
// trajectory that had failed to save.
|
||||
if len(res.Unsaved) != 1 || !strings.Contains(res.Unsaved[0], res.RunID) ||
|
||||
!strings.Contains(res.Unsaved[0], context.DeadlineExceeded.Error()) {
|
||||
t.Errorf("Unsaved = %v, want one entry naming run %s and the error", res.Unsaved, res.RunID)
|
||||
}
|
||||
}
|
||||
|
||||
// A healthy run reports nothing unsaved. Guarded so the surface never logs a
|
||||
// phantom loss.
|
||||
func TestHealthySaveReportsNothingUnsaved(t *testing.T) {
|
||||
gw := &fakeGateway{text: "the answer"}
|
||||
exec := NewModelExecutor(gw, &MemorySink{}, nil)
|
||||
res, err := exec.ExecuteAgent(context.Background(), testAgent(), testInput("q"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if len(res.Unsaved) != 0 {
|
||||
t.Errorf("Unsaved = %v on a healthy run, want none", res.Unsaved)
|
||||
}
|
||||
}
|
||||
|
||||
type failingSink struct{}
|
||||
|
||||
@@ -167,6 +167,18 @@ type ExecutionResult struct {
|
||||
// the token of whichever the person approves as ExecutionInput.Confirmation
|
||||
// on the next call.
|
||||
Confirmations []*tools.Confirmation `json:"confirmations,omitempty"`
|
||||
|
||||
// Unsaved names the trajectories this run produced that could not be
|
||||
// persisted, as "<run id>: <error>". Empty on every healthy run.
|
||||
//
|
||||
// A failed save must not fail the run — the answer already exists — but
|
||||
// it must not vanish either. Until 2026-09-22 the only record of it was
|
||||
// an entry appended to the trajectory that had just failed to save, which
|
||||
// is a note left in a bottle that sank. The runtime has no logger by
|
||||
// design; the surface does, and reads this to write the §10 line an
|
||||
// operator can grep for. Never serialised to the client: it is an
|
||||
// operator's concern, not the caller's.
|
||||
Unsaved []string `json:"-"`
|
||||
}
|
||||
|
||||
// RuntimeError is a structured error containing context for execution failures.
|
||||
|
||||
Reference in New Issue
Block a user