


The third time a payment-service deploy died on a database migration timeout, the fix was sitting in a closed incident from weeks earlier, and nobody on call could find it. That’s the whole motivation for PipelineSage: pipeline failures repeat, and the knowledge of how we fixed them last time is scattered across tickets, chat threads, and people’s heads.
I built an agent that diagnoses CI/CD failures. But the LLM call is the least interesting part. The real work went into the memory: what gets written, who is allowed to write it, and how it comes back out.
What it does and how it hangs together
PipelineSage is a small Python app with a Streamlit dashboard. You pick a failed deployment and the flow is:
failed deployment
→ recall similar incidents from Hindsight
→ re-rank the recalled memories
→ LLM diagnosis grounded only in those memories
→ recommended fix
→ human confirms the outcome
→ retain the confirmed outcome in Hindsight
The layout is deliberately boring:
app.py # Streamlit dashboard
agent/pipeline_agent.py # recall → re-rank → prompt → diagnose
memory/hindsight_memory.py # thin wrapper over Hindsight retain/recall
services/pipeline_service.py # loads pipeline runs and incident history
The model is openai/gpt-oss-120b on Groq, run at temperature=0.1. Memory lives in Hindsight Cloud. The agent never touches the Hindsight client directly; it talks to a HindsightMemory class with two methods, retain_incident and recall. That boundary turned out to matter, and I’ll come back to it.
The through-line: an agent’s memory has to be evidence, not opinion
The design question that took the most thought was what the agent is allowed to remember.
The naive version is easy: after every diagnosis, store the failure log and the model’s answer. I didn’t, because of how recall feeds back into generation. If the agent stores its own unverified recommendations, the next similar failure recalls that recommendation as “historical precedent.” Confidence compounds; correctness doesn’t.
Why Hindsight instead of my own vector table
My first instinct was a table of embeddings and a similarity query. I’ve built that before, and it works until you want anything past nearest-neighbor text matching. Then you own chunking, embedding refresh, dedup, and ranking, none of which is the thing I’m trying to build.
Hindsight, the open source agent memory system from Vectorize, gave me exactly two verbs that map onto the problem: retain to store an experience and recall to pull back what’s relevant. If the concept is new to you, Vectorize has a solid explainer on what agent memory is and how it differs from stuffing chat history into a prompt, and the Hindsight documentation covers the API I use.
The wrapper is small. This is the recall side:
python
def recall(self, query, limit=5):
result = self.client.recall(
bank_id=self.bank_id,
query=query,
max_tokens=4096,
budget=”mid”,
)
memories = []
for item in getattr(result, "results", []) or []:
text = getattr(item, "text", None)
if text:
memories.append(text)
return memories[:limit]
Enter fullscreen mode Exit fullscreen mode
What I store: structured incidents, not log dumps
Every retained incident is rendered into the same fixed shape: deployment, service, branch, environment, commit, status, failure, root cause, infrastructure change, resolution, outcome, and a pointer to the related historical incident.
python
content = f”””
DevOps pipeline incident.
Deployment: #{incident[‘deployment_id’]}
Service: {incident[‘service’]}
…
Failure:
{incident[‘error’]}
Root cause:
{incident.get(‘root_cause’, ‘Not yet confirmed.’)}
Resolution:
{incident.get(‘resolution’, ‘Not yet resolved.’)}
Outcome:
{incident.get(‘outcome’, ‘No outcome recorded.’)}
Related historical incident:
{incident.get(‘related_historical_incident’, ‘Not specified.’)}
“””
Two things fall out of this. First, the explicit Not yet confirmed. defaults mean a memory can honestly say it has no verified root cause, and the model can see that. Second, the Related historical incident field means a retained outcome links back to the precedent that produced it, so a chain like #1057 → #1017 stays legible when a human reads the memory in the dashboard.
Recall is not the end of retrieval
Semantic recall gets you candidates. It doesn’t guarantee the top result is the one you should act on. A user-service missing-environment-variable incident can be semantically “close” to a payment-service failure just because both are production deploys that failed at startup.
So the agent issues a few differently-phrased queries built from the current incident’s service and error text, de-duplicates the results, throws away the current deployment itself (an incident must never be its own precedent), and re-ranks with plain, inspectable signals:
python
if service and service.lower() in text:
score += 20 # same service
if “migration” in text:
score += 15 # same failure pattern
if “timeout” in text:
score += 10
if “succeeded” in text:
score += 25 # prefer resolutions that worked
if “batch” in text:
score += 15
if “dependency conflict” in text:
score -= 30 # different failure family
This is crude and I know it. Substring scoring is a heuristic, and I’d rather have a heuristic I can read in thirty seconds and argue with than an opaque rerank I can’t debug at 2 a.m. What matters is that successful, same-service, same-pattern incidents float to the top and unrelated families sink. The top five go into the prompt.
Constraining the model to what memory actually says
The failure mode I was most worried about wasn’t a missing answer. It was a plausible invented one: an LLM helpfully “improving” a documented fix. If memory says the fix was batches of 500 records, I do not want “batches of 200–1000 records to keep transactions under the timeout.” So the system prompt is mostly prohibitions:
- Hindsight memories are the source of truth for historical incidents.
- Never invent historical deployments, fixes, outcomes, numbers, batch sizes, timeout values, configuration values, or infrastructure changes. …
- If a historical successful resolution contains an exact value such as “500 records”, preserve that value exactly.
- If historical evidence is insufficient, clearly state that.
The response is forced into three sections: Historical Evidence, Diagnosis, Recommended Fix. That separation is what lets an engineer check the claim against the evidence directly above it.
What it looks like in use
Deployment #1057 of payment-service fails: the migration times out while updating historical transaction rows. Earlier, deployment #1017 hit a similar timeout, and the memory in Hindsight records its root cause (one large transaction exceeding the deployment timeout), its resolution (split into batches of 500 records), and its outcome (deployment succeeded).
When I run analysis on #1057, the dashboard shows the recalled memories first, then the three-part diagnosis. The recommendation is to process the historical transaction rows in batches of 500 records, citing #1017 as the evidence. After I confirm the fix worked, #1057 is retained with a link back to #1017. The next migration timeout now has two confirmed precedents instead of one.
I want to be precise about the claim. This demonstrates the loop: an outcome retained on one day changes the diagnosis on another, with the evidence on screen. I haven’t measured time-to-resolution, and I’d distrust any number I couldn’t back up.
Lessons learned
Decide who may write to memory before deciding what to store. Gating writes on human-confirmed outcomes was the most important decision in the project. An agent that remembers its own guesses builds a confident echo chamber.
Show the recalled evidence next to the answer. It turns the agent from an oracle into something an on-call engineer can audit in seconds, and it makes bad recall obvious immediately.
Retrieval needs a second pass you can read. Semantic recall gives you candidates; a small, explicit re-ranker gave me control over same-service and same-pattern preference, and excluding the incident under analysis from its own history.
Prompt against embellishment, not just against silence. The dangerous output isn’t “I don’t know.” It’s a fix with a number the memory never contained. Preserve documented values exactly and make “insufficient evidence” an acceptable answer.
What’s next
Production changes stay under human control. PipelineSage recommends; people decide. The next thing I want to capture is why a fix failed when it does, because that’s the context hardest to reconstruct later and the most valuable to recall.