Java demo repository | Difficulty: Intermediate Java | Format: Seven chapters with code and exercises

Build a Java assistant one capability at a time, following the rounds of Codepocalypse Now: LangChain4j vs. JetBrains Koog. Each chapter explains the API that introduces the next capability, gives you an experiment to run, and identifies a pattern you can use in your own application. The examples use LangChain4j builders and Java interfaces directly.

Our fictional assistant helps Viktor draft a decline for an internal training session. Calendar and organizer integrations are local MCP mocks; the organizer contacts nobody. Live chapters call hosted model APIs that may bill your accounts: Gemini for chat, Jev for decisions, Claude for drafting, and OpenAI for review. In chapters 5–7, one request can make up to seven Claude drafting calls and seven OpenAI review calls. The useful engineering problem is broader: how do you turn a generated suggestion into a grounded, reviewed, explicitly approved action?

What you’ll build

  • A Java assistant grounded in calendar tools and durable decline history.

  • Typed intent and event decisions, followed by a canonical Java request.

  • Separate Draft and Judge agents with a bounded refinement loop.

  • Human approval tied to the current message, followed by receipt-checked history.

  • A local trace showing which models, tools, and workflow stages ran.

Identify, Assemble, Draft, Judge, Human, Delivery, History. Read the top row left to right, the middle row right to left, then move down to History. Judge rejection and human feedback return to Draft. Java enforces approval of the current candidate and validates delivery receipts.

Identify selects the intent and event. Java assembles the facts, then Draft and Judge produce a reviewed proposal. Human approval and a matching delivery receipt are required before Java records the message in history. On small screens, scroll the diagram and tables sideways.

Set up the Java demo

Use JDK 21 and clone the Java demo repository. The project includes a Gradle wrapper and a launcher that downloads the pinned shared fixtures and builds the mock servers. Application and test sources are Java; the shared domain, mock servers, and TamboUI display are dependencies implemented in Kotlin.

Start from the demo project root:

git clone https://github.com/gamussa/jclaw-devoxx-be-2026.git
cd jclaw-devoxx-be-2026
cp .env.example .env
# Edit .env: set GOOGLE_API_KEY for live chat.
# Chapters 5–7 also need TYPESAFE_API_KEY, ANTHROPIC_API_KEY, and OPENAI_API_KEY.
./jclaw preview

Obtain your own Google AI Studio key for every live chapter. Chapters 5–7 also need TypeSafe, Anthropic, and OpenAI API credentials. Keep keys in your local .env, outside Git. preview displays prepared UI data without model calls or actions; it checks the layout.

The Java demo pins these LangChain4j modules:

  • 1.21.0: langchain4j, langchain4j-anthropic, langchain4j-open-ai, langchain4j-http-client-jdk.

  • 1.21.0-beta31: langchain4j-google-genai, langchain4j-agentic, langchain4j-mcp, langchain4j-typesafe.

The Gemini integration uses Google’s Java SDK, pinned to com.google.genai:google-genai:1.71.0. Gemini calls allow one attempt at both SDK and LangChain4j layers.

Keep these versions together when reproducing the workshop. The agentic and decision APIs are experimental; consult the linked documentation when upgrading.

The small Java fragments below belong inside the accompanying application and omit surrounding imports and unchanged plumbing. Names such as McpTools, SkillCatalog, Workflow, and the decline records are application classes, not LangChain4j APIs. Read the indicated source files alongside each fragment.

Java sources live under app/src/main/java/dev/gamov/jclaw/. The package names describe each class’s responsibility:

PackageClasses to explore

app

Main, DemoSession, DemoMode, SessionDisplay, CopyCommands: launch and coordinate the terminal session

agent

AgentRoles, JevDecider, GeminiDecider, TurnDecider, Workflow, WorkflowReports: typed decisions, generation and native orchestration

domain

Contracts, Delivery, SharedPolicy: wire records, receipt validation and comparison policy

model

ModelProviders, ModelLineup: provider clients and model choices

tools

McpTools, MemoryTools, SkillCatalog: scoped model capabilities

memory

SentHistory: confirmed outcome persistence

observability

TraceEvidence, GraphEvents, TraceLogProvider: listeners and local logging

serialization

Json: contract encoding and decoding

checkpoint

RoundProjection: derive the chapter checkpoints

For example, agent/Workflow.java means dev.gamov.jclaw.agent.Workflow. The typed records in the examples come from dev.gamov.jclaw.domain.Contracts. Tests mirror the production packages, with shared provider fixtures in the test-only testing package. The executable entry point is dev.gamov.jclaw.app.Main; the ./jclaw commands stay the same.

Stop the running application before changing rounds. Use the matching round/01-chatbot through round/07-observability checkpoint if your checkout includes them; otherwise, the complete main build accepts every chapter command. Make exercises on your own branch from a checkpoint so your edits remain easy to revisit. The launcher never switches branches and preserves confirmed history. Append plain to any launch command for scrolling terminal output; stop plain mode with /quit, or the dashboard with Ctrl+C.

Use the same opening request across chapters:

Get me out of the Basic AI Proficiency Training on Tuesday, run by Dana from People Ops. Don’t reuse an excuse I’ve already used on her - tell me which ones you’re avoiding.

Chapter 1 — A Java interface for conversation

Learn: ChatModel, AiServices.builder, @SystemMessage, @UserMessage, and @V.

The user enters a request in the TUI. Java AiServices binds it to the prompt and calls Gemini through ChatModel. Gemini returns text through the application. No tools or durable history are read.

Java binds the input to the AI Service prompt and calls Gemini. The model returns wording; this round reads no calendar or history.

./jclaw chatbot

The model generates text. An AI Service gives that interaction an ordinary Java method signature:

interface Assistant {
    @SystemMessage("You are Viktor's assistant. Draft suggestions only.")
    @UserMessage("{{request}}")
    String reply(@V("request") String request);
}

ChatModel model = GoogleGenAiChatModel.builder()
        .apiKey(System.getenv("GOOGLE_API_KEY"))
        .modelName(System.getenv("JCLAW_CHAT_MODEL"))
        .build();

Assistant assistant = AiServices.builder(Assistant.class)
        .chatModel(model)
        .build();

For this isolated example, set JCLAW_CHAT_MODEL to a Gemini model available to your account. The demo launcher instead resolves its role configuration and current defaults in model/ModelLineup.java. @V explicitly binds a Java argument to the prompt variable.

Try: Submit the opening request. Ask which previous excuses the assistant is avoiding. Expected: You receive text, but no tool evidence or retrieved history. A fluent answer cannot establish which excuses were actually used. Change the system message to require a shorter reply and compare the result.

Adopt: Keep the provider behind ChatModel and expose a domain-oriented interface to callers. The system prompt guides wording; later chapters add application controls for actions.

Read the code: agent/AgentRoles.java, app/DemoSession.java, model/ModelLineup.java. Official documentation: AI Services and Google Gen AI.

Chapter 2 — Ground answers with tools and MCP

Learn: @Tool, @P, .tools(…​), StdioMcpTransport, and DefaultMcpClient.

The user calls a Java AI Service, which exchanges prompts, tool requests and results with Gemini. The AI Service invokes its read-only Java tool facade. That facade uses stdio MCP to read the calendar and organizer mocks. No send tool is exposed.

Gemini proposes a tool request; the Java AI Service executes the registered facade. The facade calls the local MCP mocks and returns their results to the model.

./jclaw tools

Register a narrow Java façade over the MCP integration:

final class CalendarTools {
    private final McpTools mcp;

    CalendarTools(McpTools mcp) { this.mcp = mcp; }

    @Tool("Read the calendar. Declined entries do not contain their reasons.")
    public String getCalendar() {
        return mcp.calendarJson();
    }

    @Tool("Read sensitivity for the exact organizer.")
    public String getOrganizerSensitivity(
            @P(name = "name", description = "Exact organizer name") String name) {
        return mcp.sensitivity(name);
    }
}

Assistant assistant = AiServices.builder(Assistant.class)
        .chatModel(model)
        .tools(new CalendarTools(mcp))
        .build();

The underlying native client starts a local server process:

var transport = StdioMcpTransport.builder()
        .command(List.of(javaExecutable, "-jar", calendarJar))
        .logEvents(false)
        .build();
var client = DefaultMcpClient.builder()
        .key("calendar")
        .transport(transport)
        .build();
// The owning application closes client when the session ends.

javaExecutable and calendarJar are resolved paths in tools/McpTools.java. That class executes named MCP requests and validates results; the model-facing façade exposes only reads. MCP supplies the transport and tool contract. The registered façade determines which capabilities this assistant can request.

Try: Ask for the training time and organizer sensitivity. Inspect the tool panel or plain output. Expected: Actual calendar and organizer reads supply the answer. Previous decline reasons remain unavailable in this chapter because the calendar does not store them.

Adopt: Expose the smallest useful tool surface. Keep a future send operation under application control.

Read the code: tools/McpTools.java, especially ReadOnly and the boot method that builds each DefaultMcpClient. Official documentation: Tools and MCP integration.

Chapter 3 — Separate conversation from durable history

Learn: MessageWindowChatMemory and .chatMemory(…​); use a retrieval tool for recorded outcomes.

The Java AI Service sends context to Gemini and reads its MessageWindowChatMemory for conversation context. When Gemini requests history, Java MemoryTools reads the three seed documents and confirmed sent-history JSON. Session conversation resets on restart; outcome records remain.

The message window supplies conversation context. MemoryTools separately retrieves seed documents and confirmed sent records; calendar and organizer reads remain available.

./jclaw memory
var conversation = MessageWindowChatMemory.withMaxMessages(30);
var assistant = AiServices.builder(Assistant.class)
        .chatModel(model)
        .chatMemory(conversation)
        .tools(new McpTools.ReadOnly(mcp), memories)
        .build();

Here, memories is the application’s MemoryTools instance. The message window retains session context. It does not serve as the record of what was delivered. The separate retrieval tool reads committed prior-decline documents and confirmed literal messages in state/sent-history.json. This example uses document retrieval without an embedding store.

Try: Submit the opening request, then ask “Which reasons did you retrieve?” Restart and ask again. Expected: The three seeded reasons (calendar conflict, family obligation, and customer escalation) remain retrievable after restart. Session conversation resets; confirmed sent history persists.

Adopt: Store business outcomes separately from conversational context. For multiple users, give each conversation its own memory identity and isolate its durable records.

Read the code: tools/MemoryTools.java, memory/SentHistory.java, and the chat builder in app/DemoSession.java. Official documentation: Chat Memory and Memory in AI Services.

Chapter 4 — Load reusable instructions when needed

Learn: Runtime skill discovery and a read-only @Tool for loading a selected skill.

Java discovers the runtime skill catalog and includes its names and descriptions in the prompt sent to Gemini. Gemini can request readSkill; Java SkillCatalog validates the name and reads the selected SKILL.md from its trusted directory. Skill text returns to the model as instructions, without adding action permissions.

Java supplies catalog metadata in the system prompt. Gemini requests a selected skill body through readSkill; Java validates the selection before returning the instructions.

./jclaw skills

The demo’s SkillCatalog discovers names and descriptions from a configured directory. The assistant sees that catalog at startup and loads the relevant SKILL.md body through readSkill before applying it:

interface SkilledAssistant { String reply(String request); }

var skills = new SkillCatalog(skillRoot, toolTrace);
var assistant = AiServices.builder(SkilledAssistant.class)
        .chatModel(model)
        .systemMessageProvider(ignored ->
                "Available skills: " + skills.discover()
                + "\nRead a relevant skill before applying it. Preserve facts.")
        .tools(skills)
        .build();

skillRoot is a trusted directory and toolTrace receives evidence of reads. SkilledAssistant has no @SystemMessage; an annotated system message would take precedence over this provider. The application rejects unknown names and paths outside that root. Loading instructions does not add send permissions.

Try: Ask “Rewrite this in corporate-speak: I already build AI agents.” Then ask “Tone it down to 4.” Expected: The relevant skill is read, the first rewrite uses its default level eleven, and the follow-up reduces the style while preserving facts and commitments. Neither turn sends or saves a decline.

Adopt: Keep reusable writing rules or operating procedures outside a large permanent prompt. Load only the relevant instructions into the model context.

Framework alternative: LangChain4j also provides langchain4j-skills:1.21.0-beta31, including FileSystemSkillLoader, Skills, and a tool provider. The current demo uses its own scoped SkillCatalog; adopting the native module would be a separate implementation change.

Read the code: tools/SkillCatalog.java and the registered skills in app/DemoSession.java. Official documentation: LangChain4j Skills and Agent Skills specification.

Chapter 5 — Build a typed decision and review workflow

Learn: DecisionModel, DecisionRequest, ChoiceQuestion, typed agent methods, and native sequence, conditional, and loop builders.

The user submits a decline request. Java reads calendar and history, calls Jev for bounded intent and event choices, validates identity and assembles DeclineRequest. Its native Agentic graph calls Claude to draft DeclineDeployment, assembles the exact DeclineReview and calls OpenAI for DeclineCritique. Rejection returns to Claude for refinement. Approval ends at a reviewed proposal; round 5 has no human send gate.

Java owns identity checks, grounded reads and native graph orchestration. Jev returns choices; Claude and OpenAI return typed records. This round ends with a reviewed proposal.

./jclaw workflow

Assign a model to a specific job

StageCurrent implementationInput and output

Identify intent and event

Jev jev-1.13.0, native TypeSafe API

Original message, conversation, calendar → typed choices and probabilities

Assemble canonical request

Java application code

Verified event and organizer, retrieved memory → DeclineRequest; MCP sensitivity → OrganizerContext

Chat and skills

Gemini API

Conversation → text

Draft and Refine

Claude Opus 5.5, Anthropic API

Request and feedback → DeclineDeployment

Judge

GPT-6 Astra, OpenAI API

Current request and exact plan → DeclineCritique

Human decision and delivery

Application and native human node; local organizer mock

Approval bound to the current candidate → validated receipt

  • Jev makes the bounded intent and event choices. Java verifies the selected identities and assembles the request.

  • Gemini handles chat and skills; Claude drafts/refines; OpenAI judges. The dashboard and trace show the actual providers and IDs.

  • JCLAW_DRAFT_MODEL defaults to claude-opus-5-5; JCLAW_REVIEW_MODEL defaults to gpt-6-astra.

  • JCLAW_CHAT_MODEL chooses a Gemini model. When blank, it uses JCLAW_GEMINI_MODEL, whose default is gemini-3.7-flash.

  • Jev is the default decider. Set JCLAW_DECIDER=gemini in your environment or .env to compare Gemini’s routing; it does not act as a failure fallback.

Ask for choices rather than generated labels

This reduced example illustrates one typed routing question:

DecisionModel decider = TypeSafeDecisionModel.builder()
        .apiKey(System.getenv("TYPESAFE_API_KEY"))
        .modelName("jev-1.13.0")
        .build();

var request = DecisionRequest.builder()
        .input(userMessage)
        .question("intent", ChoiceQuestion.builder()
                .text("Does the user request a decline plan or ordinary chat?")
                .option("DECLINE", "A request to decline an obligation")
                .option("CHAT", "Conversation, explanation, or a style rewrite")
                .build())
        .build();
var intent = decider.decide(request).choice("intent");

The demo asks both intent and event questions against the same input. Event options come from the actual calendar, with NO_MATCH and AMBIGUOUS choices. Java validates the returned model identity, answer distributions, and the shared 0.60 confidence floors before assembling a canonical request. An uncertain target asks for clarification before drafting.

Give Draft and Judge different contracts

model/ModelProviders.java creates native API clients for the generation roles:

ChatModel draftModel = AnthropicChatModel.builder()
        .apiKey(System.getenv("ANTHROPIC_API_KEY"))
        .modelName("claude-opus-5-5")
        .maxTokens(8192)
        .returnThinking(false)
        .build();
ChatModel reviewModel = OpenAiChatModel.builder()
        .apiKey(System.getenv("OPENAI_API_KEY"))
        .modelName("gpt-6-astra")
        .supportedCapabilities(Capability.RESPONSE_FORMAT_JSON_SCHEMA)
        .strictJsonSchema(true)
        .store(false)
        .build();

The production builders also attach listeners, set timeouts, and disable automatic retries. returnThinking(false) selects final answer text; it does not disable Opus 5.5’s adaptive thinking. The OpenAI client requests strict JSON Schema for the typed Judge output. store(false) asks OpenAI not to store the response for later retrieval. See the Anthropic and OpenAI integration documentation.

The generation interfaces return Java records rather than free-form decisions:

interface Drafter {
    @Agent("Draft a decline proposal")
    @UserMessage("Request: {{request}}\nOrganizer: {{organizerContext}}\nPrevious: {{previous}}\nFeedback: {{feedback}}")
    DeclineDeployment draft(@V("request") DeclineRequest request,
                            @V("organizerContext") OrganizerContext organizerContext,
                            @V("previous") String previous,
                            @V("feedback") String feedback);
}

interface Critic {
    @Agent("Review the exact request and candidate")
    @UserMessage("Review: {{review}}\nOrganizer: {{organizerContext}}")
    DeclineCritique review(@V("review") DeclineReview review,
                          @V("organizerContext") OrganizerContext organizerContext);
}

DeclineReview contains both the current request and exact plan. OrganizerContext carries the exact organizer name and bounded MCP sensitivity alongside the shared records. Java checks that its organizer matches the request. Both models receive the same context through every refinement. Use the complete constraints in agent/AgentRoles.java when running the demo; these abbreviated prompts illustrate the signatures. Typed parsing establishes a shape. Java validation and a separate Judge assess whether the proposal is acceptable. Neither generation role has action tools.

One native candidate-and-review pass connects the typed roles through named state:

var draftAgent = AgenticServices.agentBuilder(Drafter.class)
        .chatModel(draftModel).outputKey("plan").build();
var assembleReview = AgenticServices.nonAiAgentBuilder(scope ->
        new DeclineReview((DeclineRequest) scope.readState("request"),
                          (DeclineDeployment) scope.readState("plan")))
        .outputKey("review").build();
var judgeAgent = AgenticServices.agentBuilder(Critic.class)
        .chatModel(reviewModel).outputKey("critique").build();

var candidatePass = AgenticServices.sequenceBuilder()
        .subAgents(draftAgent, assembleReview, judgeAgent)
        .outputKey("critique")
        .build();
var critique = (DeclineCritique) candidatePass.invoke(Map.of(
        "request", canonicalRequest,
        "organizerContext", OrganizerContext.fromMcp(
                canonicalRequest.organizerName(), "TOUCHY"),
        "previous", "No previous plan",
        "feedback", ""));

draftModel and reviewModel are independently configured ChatModel instances. canonicalRequest is assembled from verified application data before invocation. The example uses the mock’s TOUCHY result; the application passes the actual MCP response. This fragment shows one pass; the complete workflow adds validation and refinement.

Java reports the avoided reasons separately to the user. Claude writes the outward scripts, and Judge evaluates the typed candidate. Judge does not require the application’s separate report inside that candidate. Both roles must avoid reusing a reason in substance, even when the flavor label changes.

agent/Workflow.java composes those roles with AgenticServices.agentBuilder, nonAiAgentBuilder, sequenceBuilder, conditionalBuilder, and loopBuilder. Named state keys connect the stages; application nodes prepare the review and inspect its verdict. The loop allows six refinements after the initial draft: at most seven candidates per request. Exhaustion, invalid Draft or Judge output, or an unavailable Draft or Judge provider blocks the request. Automatic retries are disabled, so a single rate-limit or overload error also shows BLOCKED; start a new request after it clears. This chapter stops at a reviewed proposal.

Try: Submit the opening request, then start a separate request naming an organizer absent from the calendar. Expected: The first run shows a canonical request and typed review. The mismatched organizer requires clarification instead of silently choosing Dana. There is no human send gate in this round.

Adopt: Use a decision model for bounded choices, a generation model for drafting, and application code for identity and state validation. Carry a real refinement budget across the request.

Read the code: agent/JevDecider.java, domain/Contracts.java, agent/AgentRoles.java, and agent/Workflow.java. Official documentation: Decision Models, TypeSafe integration, Structured Outputs, and Agents and Agentic AI.

Chapter 6 — Bind human approval to the reviewed candidate

Learn: humanInTheLoopBuilder; enforce the action boundary in ordinary Java.

Starting from the assembled request, Java calls Claude Draft or Refine and OpenAI Judge. Only a Judge-approved candidate reaches Human review. Human feedback returns to Claude and fresh Judge review under the same run and shared refinement budget. Human send approval reaches the Java candidate-bound gate, which calls the organizer MCP mock. Java validates its receipt before writing confirmed sent history. Hold ends unsent, and invalid receipts leave history unchanged.

Java resumes the same request after feedback and runs fresh Refine → Judge → Human review. The application binds send approval and checks the mock receipt before writing history.

./jclaw guardrails

The native human node records a verdict supplied by the application:

var humanNode = AgenticServices.humanInTheLoopBuilder()
        .description("Human verdict on the current reviewed candidate")
        .responseProvider(scope -> scope.readState("humanFeedback"))
        .outputKey("humanVerdict")
        .build();

The terminal waits for your input between native workflow invocations. The demo does not implement durable workflow suspension or recovery after a process restart. Its human node also does not implement the application’s candidate binding or delivery validation.

Try: Inspect the reviewed email, then enter feedback before approving it. Keep the app at HUMAN; entering yes or send there approves the current message for the mock organizer. Enter this feedback:

Make the email shorter and more direct. Keep the proficiency reason; remove the Tuesday-afternoon reference.

Expected: Refine → fresh Judge → fresh human approval. The event, organizer, and run stay the same; the refinement counter continues, and Identify does not run again. Enter hold to end without sending or recording a new decline. After a reviewed candidate, send approves that exact message for the local mock organizer. Use /chat …​ for ordinary conversation while approval waits, or /new …​ to start a separate request.

The application enforces four rules:

  • Judge rejection and human feedback share one budget of six refinements.

  • A human cannot approve a candidate that Judge rejected.

  • Approval applies to the current candidate. A revised message needs fresh approval.

  • History is written only after a valid receipt matches the call, candidate, event, and organizer IDs and includes a delivery timestamp.

Delivery exercise: If you have not yet sent the proficiency-based decline, stop the app, restart ./jclaw guardrails, and submit the opening request. At a Judge-approved HUMAN candidate, enter send. Expected: DELIVERED after a matching receipt, then the literal message appears in confirmed history. If you already confirmed that decline, continue with the next fixture.

Keep confirmed history for the remaining exercises. Proficiency is now a used reason, so use this additional fictional deadline fact:

Get me out of the Basic AI Proficiency Training on Tuesday, run by Dana from People Ops. Don’t reuse a past reason. For this fictional rehearsal, I have a hard release deadline this week and am responsible for final release verification. Ask Dana for permission to prioritize that work over the training. Use only the deadline reason; don’t add my conference talk, proficiency, or a calendar conflict. Tell me which past reasons you’re avoiding.

Restarting clears current-session proposed alternatives while retaining confirmed history.

Failure exercise: Stop the app and launch JCLAW_MOCK_DELIVERY=wrong-candidate ./jclaw guardrails plain. Submit the deadline request above, then enter send at a Judge-approved HUMAN candidate. Expected: Delivery stays unconfirmed and history is not updated. There is no automatic resend.

Adopt: Treat approval as a decision about specific content. Record only confirmed outcomes, and reconcile uncertain delivery before retrying. A content hash binds a candidate; it does not provide authentication or server-side idempotency.

Read the code: agent/Workflow.java (ApprovalGate), app/DemoSession.java, domain/Delivery.java, and memory/SentHistory.java. Official documentation: Human in the loop. This chapter’s delivery gate is application logic; LangChain4j input and output guardrails are a separate API.

Chapter 7 — Observe what actually ran

Learn: ChatModelListener, DecisionModelListener, native AgentListener callbacks, and HTML reports from AgentMonitor.

Native model listeners, Agentic callbacks, and separate Java and MCP events feed TraceEvidence for local JSONL and TUI views. AgentMonitor separately observes the native review-loop and human roots, and HtmlReportGenerator writes their topology and execution timelines to local HTML files. Jev Identify, ordinary chat, MCP reads, delivery and history stay in JSONL and TUI evidence. No hosted exporter is configured.

Native model and graph listeners emit their own observations. Java and MCP events describe application operations; TraceEvidence correlates them for the local file and dashboard.

./jclaw observability

A minimal chat listener can measure a model interaction without logging its prompt:

ChatModelListener listener = new ChatModelListener() {
    private final Object started = new Object();

    @Override
    public void onRequest(ChatModelRequestContext context) {
        context.attributes().put(started, System.nanoTime());
    }

    @Override
    public void onResponse(ChatModelResponseContext context) {
        long elapsed = System.nanoTime()
                - (Long) context.attributes().get(started);
        System.out.println("durationMs=" + elapsed / 1_000_000
                + " usage=" + context.chatResponse().tokenUsage());
    }
};
// Attach with .listeners(List.of(listener)) on each provider's model builder.

The complete demo also records model errors, Jev’s typed decisions and probabilities, and native agent callbacks. Application events separately describe MCP reads, canonical request assembly, human waiting, delivery receipts, and history writes. The local trace is state/trace.jsonl; it includes fictional model inputs and outputs and should not be treated as a production logging policy.

LangChain4j can also generate a self-contained HTML report of a native graph’s topology and execution history. Attach a monitor to the sequence from Chapter 5:

AgentMonitor monitor = new AgentMonitor();
var monitoredPass = AgenticServices.sequenceBuilder()
        .name("draftAndJudge")
        .subAgents(draftAgent, assembleReview, judgeAgent)
        .listener(monitor)
        .build();
monitoredPass.invoke(Map.of(
        "request", canonicalRequest,
        "organizerContext", OrganizerContext.fromMcp(
                canonicalRequest.organizerName(), "TOUCHY"),
        "previous", "No previous plan",
        "feedback", ""));
HtmlReportGenerator.generateReport(monitor, Path.of("review-loop.html"));

This snippet illustrates report wiring; the demo’s full workflow adds the bounded loop, safety checks, and shared human-refinement budget. The app attaches one monitor to the review loop and another to the human sequence, then updates both reports after each turn. Enter /report in the TUI or plain app to get their paths: state/reports/<traceId>/review-loop.html and human-review.html. Open the files in a browser and refresh after a revision or verdict.

The native reports show Draft/Judge iterations and submitted human verdicts. Jev Identify, ordinary chat, MCP reads, approval gating, delivery and history remain in JSONL/TUI evidence. Native scope/session IDs differ from the application trace ID in the directory name. A native "success" means the invocation completed; it does not mean the draft was approved or delivered.

Try: Use the fictional deadline request from Chapter 6, give human feedback once, inspect the revised review, and choose hold. Use F3 TRACE to inspect the actual inputs, outputs, usage, and native callbacks. Use /report to open the review topology and timeline, then the human verdict report after hold. The file appends across launches, and traceId identifies an application session. After stopping, count events from the last session with the optional jq command:

jclaw_trace_id=$(tail -n 1 state/trace.jsonl | jq -r '.traceId')
jq -n --arg traceId "$jclaw_trace_id" '
  reduce inputs as $event (
    {DECISION_OUTPUT: 0, MODEL_OUTPUT: 0, AGENTIC_GRAPH: 0,
     CANDIDATE: 0, VERDICT: 0, HUMAN_HOLD: 0,
     DELIVERED: 0, MEMORY_SAVED: 0};
    if $event.traceId == $traceId then
      if has($event.kind) then .[$event.kind] += 1 else . end
    else . end
  )
' state/trace.jsonl

Expected: One Identify pass for the decline request, evidence for each candidate and Judge interaction, a human verdict, and no delivery/history-write event for the held candidate. The review HTML shows both native invocations, and the human HTML records feedback and hold. Separate native graph callbacks from application stages when explaining the trace. API usage and elapsed time are visible; there is no hosted telemetry exporter or subscription CLI usage in this implementation.

Adopt: Correlate model, tool, workflow, and business events. Attach a role and actual provider to measurements, and choose which content your production logs may retain.

For a provider-free example, run ./gradlew :app:reportFixture --offline and open state/report-fixture/review-loop.html and human-review.html. These are prepared-response fixtures of real native execution, not reports from your live session.

Read the code: observability/TraceEvidence.java, observability/GraphEvents.java, agent/WorkflowReports.java, and report wiring in app/Main.java. Official documentation: Observability, Decision model listeners, and Agentic monitoring and HTML reports.

Continue in your own application

Choose a workflow with a concrete reviewable result, such as a support-ticket reply, a change proposal, or a calendar suggestion. Keep the Java records, limited tool surface, separate outcome history, bounded review loop, and approval of exact content. Replace the scenario and integrations one at a time, verifying the contract at each boundary.

For the accompanying demo, the provider-free validation command is:

./gradlew :app:check :app:installDist :mocks:mcpJars --console=plain

The checks exercise typed parsing, workflow limits, approval binding, delivery failure cases, and restart persistence. They use network-boundary model fixtures; a passing test is separate from a live provider result.

For more context, read LangChain4j for Java developers or explore the Kafka, Flink, and Tableflow workshop to practice grounding applications in real data. Viktor’s related talks include Codepocalypse Now and MCP integrations.

About the author

Viktor Gamov is a Java Champion, developer advocate, and coauthor of Kafka in Action. He teaches Java, data systems, and AI application development through conference talks and hands-on workshops. See his background, GitHub projects, and speaking history.

References

Showdown provider note: The comparison handoff assigns Claude (Anthropic) subscription CLI to Draft/Refine and Codex (OpenAI) subscription CLI to Judge. This Java version uses Claude Opus 5.5 through the Anthropic API and GPT-6 Astra through the OpenAI API. Call its critic an OpenAI API Judge; it does not launch Codex CLI. Compare API and subscription usage separately. The selected IDs are documented in Claude Opus 5.5 and GPT-6 Astra. These are runtime model roles, separate from the coding assistants used to build the demo. The Koog app wires those calls in Strategy.kt, constructs Claude agents in CliCritic.kt, and launches codex exec --output-schema through TypedCodex.kt.