Context growing unboundedly: 517 → 2101 tokens across 3 calls
Input token counts are climbing with each LLM call in this trace (517 → 1844 → 2101). This is the unbounded conversation history pattern — cost and latency grow O(N²) as the session continues. At this rate, long sessions will become expensive and slow.
Fix: trim context with a rolling window or summarisation
# Option 1: rolling window — keep only the last N messages
MAX_MESSAGES = 10
messages = messages[-MAX_MESSAGES:]
# Option 2: summarise when context exceeds a threshold
def trim_context(messages, max_tokens=3000):
total = count_tokens(messages)
if total <= max_tokens:
return messages
# Summarise the oldest half, keep the recent half verbatim
old = messages[:-5]
recent = messages[-5:]
summary = llm.summarise(old)
return [{"role": "system", "content": f"Earlier context: {summary}"}] + recentWaterfall
workflow.execute
anthropic.messages.create
Summarize the Q3 earnings report attached. Focus on revenue, margin, and guidance.
(answer to: Summarize the Q3 earnings report attache…)
anthropic.messages.create
Given the customer ticket below, draft a refund response that follows policy P-204.
(answer to: Given the customer ticket below, draft a…)
anthropic.messages.create
Translate the user manual section to Japanese, preserving the table structure.
(answer to: Translate the user manual section to Jap…)
Faithfulness · anthropic.messages.create0.25
1 supported · 1 contradicted · 2 unsupported of 4 claims
- supported
Day 1 plan fits under the $200/day soft cap.
- contradicted
Belém Tower opens at 9am.
- unsupported
Time Out Market lunch costs $18.
- unsupported
Castelo dinner costs $42.