Implementing LLM Context Summarization in Java: A Production-Grade Design

Long-running LLM conversations eventually collide with a hard constraint: the model’s context window. Even before the hard limit is reached, repeatedly sending the entire conversation increases latency and cost while giving the model more irrelevant history to sift through.

A common solution is context summarization: replace older conversation history with a compact summary while preserving the most recent messages exactly. The idea sounds simple, but a production implementation must handle system instructions, token thresholds, tool calls, persistent memory, concurrent requests, streaming, and failures without corrupting the conversation.

This article develops a generic Java design for context summarization middleware. It does not depend on a particular web framework or LLM provider.

1. The desired transformation

Suppose the current conversation is:

System: You are a travel assistant.
User: I need a flight to Delhi.
Assistant: What date?
User: October 20.
Assistant: Morning or evening?
User: Morning.

If the policy says to retain the two most recent messages, the compacted conversation becomes:

System: You are a travel assistant.
User: Here is a summary of the conversation so far:

      The user needs a flight to Delhi on October 20.
Assistant: Morning or evening?
User: Morning.

Three properties are important:

  • The original system message remains a real system message.
  • Old messages are replaced, not merely followed by a summary.
  • Recent messages remain verbatim because summarization is lossy.

2. Separate triggering from retention

Two independent policies are required:

  • Trigger: when should summarization run?
  • Retention: how much recent history should remain unchanged?

A useful configuration model is:

record SummaryTrigger(Integer messages, Integer tokens) {}
record Retention(Integer messages, Integer tokens) {}

The trigger can use a message threshold, token threshold, or both:

trigger = new SummaryTrigger(20, null);       // 20 messages
trigger = new SummaryTrigger(null, 8_000);    // 8,000 tokens
trigger = new SummaryTrigger(20, 8_000);      // both must be reached

Using AND semantics when both values are present makes the configuration predictable:

boolean shouldSummarize(List<ChatMessage> messages) {
    boolean messagesReached =
        trigger.messages() == null
            || messages.size() >= trigger.messages();

    boolean tokensReached =
        trigger.tokens() == null
            || estimator.estimate(messages) >= trigger.tokens();

    return messagesReached && tokensReached;
}

Retention should use exactly one mode:

new Retention(6, null);       // keep six recent messages
new Retention(null, 2_000);   // keep complete recent messages within 2,000 tokens

3. Calculate the preferred split

Consider ten messages and keep.messages = 3:

0 1 2 3 4 5 6 | 7 8 9
summarize         retain

The preferred cutoff is:

candidate = messages.size() - keep.messages()
          = 10 - 3
          = 7

If index zero is a system message, the cutoff must never move before index one:

int candidate = Math.max(historyStart, messages.size() - keepMessages);

Token retention walks backward from the newest message and keeps complete messages while they fit:

int tokenRetentionCutoff(
        List<ChatMessage> messages,
        int historyStart,
        int keepTokens) {

    int retainedTokens = 0;
    int cutoff = messages.size();

    for (int i = messages.size() - 1; i >= historyStart; i--) {
        int messageTokens = Math.max(0, estimator.estimate(messages.get(i)));

        if (cutoff < messages.size()
                && retainedTokens + messageTokens > keepTokens) {
            break;
        }

        retainedTokens += messageTokens;
        cutoff = i;
    }

    return cutoff;
}

This keeps whole messages. Splitting a message by characters or tokens can damage structured content and tool metadata.

4. Never split a tool request from its result

Agent conversations contain more than user and assistant text. An assistant can request one or more tools, followed by tool-result messages:

index 0  User: Check Delhi and Mumbai weather
index 1  Assistant:
           ToolCall(id=call-1, city=Delhi)
           ToolCall(id=call-2, city=Mumbai)
-------- candidate cutoff = 2 --------
index 2  ToolResult(id=call-1, text=20 C)
index 3  ToolResult(id=call-2, text=27 C)
index 4  User: Which city is cooler?

Keeping indexes 2 through 4 would retain tool results while removing the assistant message that requested those tools. Many model APIs reject this history; even when accepted, it is semantically incomplete.

The cutoff must move backward to index one:

summarize: [0]
retain:    [1, 2, 3, 4]

A safe-boundary algorithm collects consecutive tool-result IDs and searches backward for the matching assistant request:

static int safeToolBoundary(
        List<ChatMessage> messages,
        int historyStart,
        int candidate) {

    if (candidate >= messages.size()
            || !(messages.get(candidate) instanceof ToolResult)) {
        return candidate;
    }

    Set<String> requestIds = new HashSet<>();
    int afterToolResults = candidate;

    while (afterToolResults < messages.size()
            && messages.get(afterToolResults) instanceof ToolResult result) {
        if (result.requestId() != null) {
            requestIds.add(result.requestId());
        }
        afterToolResults++;
    }

    for (int i = candidate - 1; i >= historyStart; i--) {
        if (messages.get(i) instanceof AssistantMessage assistant
                && assistant.toolCalls().stream()
                    .map(ToolCall::id)
                    .anyMatch(requestIds::contains)) {
            return i;
        }
    }

    return afterToolResults;
}

If no matching request exists, those results are already orphaned. Moving the cutoff forward past the result block avoids retaining invalid recent history:

index 0  User: old conversation
index 1  Assistant: old answer
-------- candidate = 2 --------
index 2  ToolResult(id=unknown-1)
index 3  ToolResult(id=unknown-2)
index 4  User: continue

summarize: [0, 1, 2, 3]
retain:    [4]

5. Why format the old messages?

The old conversation is not one string. It is a typed list containing user messages, assistant messages, system messages, tool calls, tool results, and potentially multimodal content.

The summarizer should receive one instruction plus one user message containing the history as data:

SystemMessage: "Summarize the important context..."
UserMessage:   "<user>...</user><assistant>...</assistant>"

Formatting converts typed message objects into that single role-preserving string:

<user>Check the weather in Delhi.</user>
<assistant>
  <tool-call id="weather-1" name="getWeather">
    {"city":"Delhi"}
  </tool-call>
</assistant>
<tool-result id="weather-1" name="getWeather">
  20 C
</tool-result>
<assistant>It is 20 C in Delhi.</assistant>

The XML-like tags are just structured text. They make role boundaries and tool relationships explicit. Simply joining message text would lose who said what and could turn tool arguments or results into ambiguous prose.

Escape text before inserting it into this representation. For non-text content, use a safe placeholder such as [image content] rather than copying a large base64 payload into the summarization request.

6. Use a raw model for the summary call

The summary call should bypass the agent’s middleware pipeline:

ChatResponse generateSummary(String formattedHistory) {
    ChatRequest request = ChatRequest.of(
        SystemMessage.from(summaryPrompt),
        UserMessage.from(formattedHistory)
    );

    ChatResponse response = rawSummaryModel.chat(request);

    if (response == null
            || response.text() == null
            || response.text().isBlank()) {
        throw new InvalidSummaryException();
    }

    return response.text().trim();
}

Calling the already-wrapped agent model can recursively invoke summarization. It may also incorrectly consume another logical model-call allowance or apply fallback and retry policies intended for the user’s primary request.

A clean policy is:

  • Use an explicitly configured summarizer model when supplied.
  • Otherwise use the raw primary model.
  • Make one summary attempt.
  • Propagate failures; do not fabricate a summary.

7. Reconstruct both the request and persistent memory

After obtaining a summary, build a new list:

List<ChatMessage> compacted = new ArrayList<>();
compacted.addAll(messages.subList(0, historyStart));
compacted.add(UserMessage.from(summaryPrefix + "\n\n" + summary));
compacted.addAll(messages.subList(cutoff, messages.size()));

The summary is conversation context, not a privileged instruction, so representing it as a user message avoids giving generated prose system-level authority.

Both surfaces must change:

  • The current request must use the compacted list immediately.
  • Persistent chat memory must store the compacted list for future calls.

Updating only the request means the next call reloads the original long memory. Updating only memory means the current provider call can still receive the original long request.

8. Make memory replacement concurrency-safe

Summary generation is a network call. Another request can modify the same conversation while it is in progress:

Thread A: reads messages [0..10]
Thread A: starts summary model call
Thread B: appends message 11
Thread A: replaces memory using its stale [0..10] snapshot
Result:   message 11 is lost

Use compare-before-replace semantics:

synchronized boolean replaceIfUnchanged(
        List<ChatMessage> expected,
        List<ChatMessage> replacement) {

    List<ChatMessage> current = List.copyOf(delegate.messages());
    if (!current.equals(expected)) {
        return false;
    }

    delegate.replaceAll(replacement);
    return true;
}

Then the middleware behaves safely:

if (!memory.replaceIfUnchanged(original, compacted)) {
    return originalRequest;
}

return originalRequest.withMessages(compacted);

A stale summary call may be wasted and one oversized request may proceed, but no conversation data is lost. Distributed deployments require the same compare-and-swap guarantee in the shared memory store; a JVM synchronized block protects only one process.

9. Put logical model-call limits outside summarization

A useful middleware order is:

model-call limit
  -> summarization
    -> model fallback
      -> model retry
        -> physical provider

This ordering has deliberate semantics:

  • If the logical call limit is already exhausted, the summarizer is not invoked.
  • The internal summary call is not counted as a second user-visible model call.
  • If summarization fails, the outer reservation is released.
  • Fallback and retry apply to the primary response, not to summary generation.

10. Streaming does not make summarization asynchronous

Summarization must finish before the streaming provider receives its request. Therefore no partial response should be visible before compaction succeeds:

try {
    ChatRequest compacted = summarizer.beforeModel(request);
    streamingModel.chat(compacted, handler);
} catch (RuntimeException failure) {
    handler.onError(failure);
}

Once provider streaming begins, changing its conversation context is too late. A summary failure should follow the normal streaming error path and produce no partial provider output.

11. Failure behavior should be atomic

Do not mutate memory before the summary model succeeds. The correct sequence is:

  1. Read a snapshot.
  2. Calculate and validate the cutoff.
  3. Format the old history.
  4. Call the summary model.
  5. Reject null or blank summaries.
  6. Construct the replacement list.
  7. Compare and replace memory.
  8. Return the rebuilt request.

If steps three through five fail, the original memory remains untouched.

12. A compact middleware skeleton

ChatRequest beforeModel(
        ExecutionContext execution,
        ChatRequest request) {

    ConversationMemory memory = execution.memory();
    if (memory == null) {
        throw new MemoryRequiredException();
    }

    List<ChatMessage> messages = List.copyOf(request.messages());
    if (!shouldSummarize(messages)) {
        return request;
    }

    int historyStart = hasLeadingSystemMessage(messages) ? 1 : 0;
    int cutoff = calculateRetentionCutoff(messages, historyStart);
    cutoff = safeToolBoundary(messages, historyStart, cutoff);

    if (cutoff <= historyStart) {
        return request;
    }

    List<ChatMessage> oldHistory =
        messages.subList(historyStart, cutoff);

    List<ChatMessage> trimmed =
        trimForSummaryModel(oldHistory);

    String formattedHistory = formatMessages(trimmed);
    String summary = generateSummary(formattedHistory);

    List<ChatMessage> compacted = new ArrayList<>();
    compacted.addAll(messages.subList(0, historyStart));
    compacted.add(UserMessage.from(summaryPrefix + "\n\n" + summary));
    compacted.addAll(messages.subList(cutoff, messages.size()));

    if (!memory.replaceIfUnchanged(messages, compacted)) {
        return request;
    }

    return request.withMessages(compacted);
}

13. Tests that matter

Focused tests should cover:

  • Message-only, token-only, and combined trigger thresholds.
  • Message- and token-based retention.
  • Leading system-message preservation.
  • Parallel tool calls and consecutive tool results.
  • Orphaned tool results.
  • Role and tool metadata in formatted history.
  • Empty and failed summary responses.
  • Memory remaining unchanged after failure.
  • Concurrent modification rejecting stale replacement.
  • Per-conversation memory isolation.
  • Model-call limits running before summarization.
  • Streaming errors occurring before any partial response.

Conclusion

Reliable context summarization is not merely “ask the model for a summary.” It is a transactional transformation of structured conversation history.

A robust Java implementation:

  • separates trigger and retention policy;
  • preserves system instructions and recent messages;
  • keeps tool requests and results structurally valid;
  • formats typed history as explicit role-preserving data;
  • uses a raw model without recursive middleware;
  • updates the current request and persistent memory together;
  • rejects stale concurrent replacement;
  • and fails without damaging conversation state.

With those guarantees in place, summarization becomes a dependable context-management mechanism rather than a best-effort prompt trick.

You may also like...

Leave a Reply

Your email address will not be published. Required fields are marked *