Spring Boot 3, Virtual Threads (Project Loom), and Stripe Integration: How @Retryable on Virtual-Thread-Backed Services Re-invokes the Method Body, newVirtualThreadPerTaskExecutor() Manual Retry Loops Regenerate Keys, and StructuredTaskScope Fork Restart Creates New Charges
Java 21’s virtual threads let you write blocking code at scale — but blocking code with retry logic and UUID.randomUUID() for Stripe idempotency keys is exactly where Project Loom’s greatest misconception lives. Virtual threads change which carrier thread executes your code; they do not change when your code executes, how many times it executes, or what state is preserved between those executions.
This post covers three Spring Boot 3 + virtual threads + Stripe integration failure modes that are structurally distinct from the basic Spring Boot @Transactional post, the Spring @Async + @Transactional post, and the CompletableFuture batch billing post. The modes here are specific to the virtual-thread execution model introduced in Java 21 and stabilized in Spring Boot 3.2: how spring.threads.virtual.enabled=true interacts with Spring Retry’s AOP model, how the cheap blocking semantics of virtual threads encourage a specific manual retry pattern that generates duplicate charges, and how Java 21’s structured concurrency API (StructuredTaskScope) creates a new failure mode when combined with an outer retry loop.
Background: virtual threads, continuations, and the Stripe idempotency contract
Virtual threads (JEP 444, finalized in Java 21) are user-mode threads managed by the JVM rather than the operating system. Each virtual thread has its own call stack and local variable frame, just like a platform thread. The key difference is mounting: a virtual thread is mounted onto a carrier (platform) thread only when it is runnable. When a virtual thread reaches a blocking operation — Thread.sleep(), socket I/O, Object.wait() — it unmounts from the carrier thread and suspends its continuation (the saved call stack). The carrier thread is free to run other virtual threads. When the blocking operation completes, the virtual thread is rescheduled onto a carrier thread and its continuation is resumed.
The continuation mechanism is an implementation detail of how virtual threads handle blocking. From the perspective of the code running inside a virtual thread, execution is sequential: the call stack advances one frame at a time. A method call inside a virtual thread is exactly as sequential as a method call inside a platform thread. There is no “snapshot” of local variable state that persists between separate method calls. Each call to a method — whether a first invocation or a retry invocation via AOP — creates a fresh stack frame with fresh local variables.
Spring Boot 3.2 introduced spring.threads.virtual.enabled=true, which configures Tomcat’s connector thread pool, @Async executors, and scheduled task executors to use virtual threads. This is a significant operational change — it eliminates the need to size thread pools for I/O-bound workloads. It does not change any application-level semantics: Spring Retry’s @Retryable AOP advice, transaction propagation, bean scope, or any other Spring framework behavior.
Stripe’s idempotency key contract: Stripe deduplicates requests based on the Idempotency-Key header. Two requests with the same key for the same endpoint within 24 hours return the cached result of the first successful request. Two requests with different keys for the same customer at the same amount are two charges. The contract is per-key, per-endpoint, not per-customer, not per-amount.
Mode 1: @Retryable on a virtual-thread-backed service — UUID inside the service method generates UUID_B on the retry invocation — ch_B alongside committed ch_A
When spring.threads.virtual.enabled=true is set in Spring Boot 3.2+, the executors backing @Async methods and Tomcat connectors are replaced with virtual-thread executors. A @Retryable annotation on a service method runs on whatever thread calls that method — which, with virtual threads enabled, may be a virtual thread. This changes nothing about @Retryable’s retry semantics.
Spring Retry’s @Retryable is implemented as AOP advice. The generated proxy intercepts the method call and wraps it in a RetryTemplate. On a retryable exception, RetryTemplate calls proceed() on the method interceptor again. proceed() is a fresh invocation of the intercepted method. The JVM creates a new stack frame for the method body. Every local variable inside the method is initialized from scratch. UUID.randomUUID() inside the method body generates a new UUID on every invocation — UUID_B on the second attempt.
// BillingService.java — unsafe mode 1 (UUID inside @Retryable method)
@Service
public class BillingService {
private final StripeClient stripeClient;
private final BillingRecordRepository repo;
// Developer's reasoning:
// "spring.threads.virtual.enabled=true means this runs on a virtual thread.
// Virtual threads use continuations under the hood — like Kotlin coroutines.
// When @Retryable retries, maybe it resumes the continuation rather than
// starting a new invocation. The UUID was already computed on attempt 1;
// if the continuation is resumed, the UUID binding might be preserved."
//
// The problem: @Retryable calls proceed() — a fresh invocation of charge().
// Virtual threads are a scheduling mechanism. The continuation governs
// parking/unparking at I/O and sleep boundaries, not method call semantics.
// proceed() creates a new stack frame regardless of thread type.
// UUID.randomUUID() inside the method body generates UUID_B on the retry.
// Attempt 1: UUID_A sent. Stripe commits ch_A. StripeConnectException thrown.
// Attempt 2: proceed() invoked. New stack frame. UUID_B generated.
// Stripe sees a new key — commits ch_B. Two charges, one customer.
@Retryable(
retryFor = StripeConnectException.class,
maxAttempts = 3,
backoff = @Backoff(delay = 500, multiplier = 2)
)
public PaymentIntent charge(String customerId, long amountCents, String billingPeriod)
throws StripeException {
// UUID computed here — fresh on every proceed() invocation
String idempotencyKey = UUID.randomUUID().toString(); // UUID_B on retry
PaymentIntentCreateParams params = PaymentIntentCreateParams.builder()
.setAmount(amountCents)
.setCurrency("usd")
.setCustomer(customerId)
.setConfirm(true)
.build();
PaymentIntent pi = stripeClient.paymentIntents().create(
params,
RequestOptions.builder().setIdempotencyKey(idempotencyKey).build()
);
repo.save(new BillingRecord(customerId, pi.getId(), billingPeriod, amountCents));
return pi;
}
}
# application.properties — Spring Boot 3.2 virtual threads
spring.threads.virtual.enabled=true
The developer’s mental model is wrong at a fundamental level. Kotlin coroutines and JVM virtual threads share the word “continuation” in their implementation descriptions, but they mean different things in different contexts. A Kotlin coroutine’s continuation captures the local variable state at a suspension point and restores it when the coroutine is resumed. A virtual thread’s continuation captures the call stack at a blocking point so the carrier thread can be released. Neither mechanism persists local variable state across separate method calls initiated by an external caller (such as proceed() from an AOP interceptor).
Why the misconception is particularly sticky for virtual threads
The confusion arises from the legitimate analogy between virtual threads and coroutines at the scheduling level. Both allow blocking code to be written without blocking an OS thread. Both use a form of continuation to save and restore call stack state. A developer who understands Kotlin coroutines and has read that virtual threads are “JVM-native coroutines” (a common informal description) may reasonably extrapolate that @Retryable on a virtual thread works like retry in a Kotlin suspend function — where a suspension-aware retry framework (kotlinx-retry, or a custom Flow.retry{}) can in principle resume the coroutine from a saved state.
But even in Kotlin, this is not how coroutine-aware retry works. Kotlin coroutine retry re-invokes the suspend function body from the first statement, not from a suspension point mid-function. And @Retryable is not a coroutine-aware retry framework — it is a Java AOP framework that calls proceed(), which always starts the method from its first statement regardless of whether it runs on a virtual thread, a coroutine, a platform thread, or any other execution model.
The fix: content-hash key computed before the retryable call
The key must be computed before the @Retryable proxy boundary. If the key is computed in the caller and passed as a parameter to the service method, the same key is used on every retry invocation of charge().
// BillingService.java — fixed mode 1 (key passed as parameter)
@Service
public class BillingService {
private final StripeClient stripeClient;
private final BillingRecordRepository repo;
@Retryable(
retryFor = StripeConnectException.class,
maxAttempts = 3,
backoff = @Backoff(delay = 500, multiplier = 2)
)
public PaymentIntent charge(String customerId, long amountCents,
String billingPeriod, String idempotencyKey)
throws StripeException {
// Key is a parameter — stable across all proceed() invocations.
// Attempt 1: key passed in → Stripe commits ch_A → exception.
// Attempt 2: same key passed in → Stripe deduplicates → returns ch_A.
// No ch_B.
PaymentIntentCreateParams params = PaymentIntentCreateParams.builder()
.setAmount(amountCents)
.setCurrency("usd")
.setCustomer(customerId)
.setConfirm(true)
.build();
PaymentIntent pi = stripeClient.paymentIntents().create(
params,
RequestOptions.builder().setIdempotencyKey(idempotencyKey).build()
);
repo.save(new BillingRecord(customerId, pi.getId(), billingPeriod, amountCents));
return pi;
}
}
// BillingOrchestrator.java — key computed in the caller before @Retryable boundary
@Service
public class BillingOrchestrator {
private final BillingService billingService;
public void billCustomer(String customerId, long amountCents, String billingPeriod)
throws StripeException {
// Content-hash key computed here, outside the @Retryable boundary.
// This method is not @Retryable — it calls a @Retryable method.
// The key is stable: same customer + period + amount = same key, every time.
String idempotencyKey = contentHashKey(customerId, billingPeriod, amountCents);
billingService.charge(customerId, amountCents, billingPeriod, idempotencyKey);
}
private String contentHashKey(String customerId, String period, long amountCents) {
String input = customerId + "|" + period + "|" + amountCents;
try {
MessageDigest md = MessageDigest.getInstance("SHA-256");
byte[] hash = md.digest(input.getBytes(StandardCharsets.UTF_8));
StringBuilder sb = new StringBuilder();
for (byte b : hash) sb.append(String.format("%02x", b));
return sb.toString();
} catch (NoSuchAlgorithmException e) {
throw new IllegalStateException("SHA-256 not available", e);
}
}
}
The alternative is to move the key computation into the @Retryable method itself, but before any statement that could throw. A key computed at line 1 of the method is still regenerated on retry because line 1 runs on every proceed() invocation. The only safe placement for the key is outside the retry boundary: either as a parameter, or in a non-retryable wrapper that computes and caches the key before calling the retryable inner method.
Virtual threads and the @Backoff delay: a genuine benefit
One area where virtual threads do genuinely improve the @Retryable experience is the backoff delay. @Retryable’s @Backoff implementation calls Thread.sleep() between retry attempts. On a platform thread, Thread.sleep(500) blocks the thread for 500 ms, holding a thread-pool thread idle. With spring.threads.virtual.enabled=true, the sleeping thread is a virtual thread. It unmounts from the carrier thread during the sleep, allowing the carrier to serve other requests. A service with 100 concurrent @Retryable calls in their backoff sleep no longer needs 100 platform threads parked. This is a real operational benefit — it just does not affect idempotency.
Test: verifying the key is stable across retry attempts
// BillingServiceRetryTest.java — JUnit 5 + WireMock
@SpringBootTest
@TestPropertySource(properties = "spring.threads.virtual.enabled=true")
class BillingServiceRetryTest {
@Autowired
BillingOrchestrator orchestrator;
@RegisterExtension
static WireMockExtension wireMock = WireMockExtension.newInstance()
.options(wireMockConfig().dynamicPort())
.build();
@Test
void retryUsesTheSameIdempotencyKey() throws Exception {
// Scenario: first call returns 500, second call returns 200.
// Assert: both calls use the same Idempotency-Key header.
wireMock.stubFor(post(urlPathEqualTo("/v1/payment_intents"))
.inScenario("retry").whenScenarioStateIs(STARTED)
.willReturn(serverError())
.willSetStateTo("attempt_2"));
wireMock.stubFor(post(urlPathEqualTo("/v1/payment_intents"))
.inScenario("retry").whenScenarioStateIs("attempt_2")
.willReturn(ok().withBody(paymentIntentJson("pi_test_1"))));
orchestrator.billCustomer("cus_A", 5000L, "2026-10");
List<LoggedRequest> requests = wireMock.findAll(
postRequestedFor(urlPathEqualTo("/v1/payment_intents")));
assertThat(requests).hasSize(2);
String key1 = requests.get(0).getHeader("Idempotency-Key");
String key2 = requests.get(1).getHeader("Idempotency-Key");
assertThat(key1).isNotBlank();
assertThat(key2).isEqualTo(key1); // same key on both attempts
}
}
If the implementation has the UUID inside the service method (unfixed), this test fails: key2 is a different UUID from key1. The test failure surface is clear: two different Idempotency-Key header values across attempts is definitive proof of a duplicate-charge risk.
Mode 2: Executors.newVirtualThreadPerTaskExecutor() + manual Thread.sleep() retry loop — UUID inside the loop body generates UUID_B on each iteration — ch_B per customer
Virtual threads make Thread.sleep() cheap. A virtual thread sleeping for 500 ms unmounts from its carrier and parks, consuming no OS resources beyond the virtual thread’s stack. This is the primary reason the reactive vs blocking debate changes when virtual threads are available: the traditional argument for reactive code (“blocking is wasteful”) no longer applies to I/O-bound work on virtual threads.
A common migration pattern from reactive retry operators to virtual-thread code is to replace Mono.retryWhen(Retry.backoff(3, Duration.ofMillis(500))) with a manual while loop and Thread.sleep(). The developer sees this as a simplification: the reactive retry operator had complex configuration; the loop is plain Java. The problem is that UUID.randomUUID() inside a loop body is a new local variable binding on every iteration.
// BillingWorker.java — unsafe mode 2 (UUID inside virtual-thread retry loop)
@Component
public class BillingWorker {
private static final ExecutorService EXECUTOR =
Executors.newVirtualThreadPerTaskExecutor();
private final StripeClient stripeClient;
private final BillingRecordRepository repo;
public CompletableFuture<PaymentIntent> chargeAsync(
String customerId, long amountCents, String billingPeriod) {
return CompletableFuture.supplyAsync(() -> {
int attempt = 0;
StripeException lastException = null;
while (attempt < 3) {
try {
// Developer's reasoning:
// "UUID is inside the loop but at the top, before the Stripe call.
// Each attempt gets a fresh UUID — that's intentional, because
// if the first attempt committed, a second attempt with the same key
// would just return the cached charge (idempotency works).
// If the first attempt timed out, we don't know if it committed,
// so a new UUID ensures we get a new charge attempt recorded.
// Virtual threads mean the sleep is cheap — no thread waste."
//
// The problem: Stripe's idempotency key is the ONLY mechanism for
// Stripe to deduplicate charges. A new UUID on attempt 2 means
// Stripe has no way to know ch_A was committed on attempt 1.
// If the StripeConnectException on attempt 1 was thrown AFTER
// Stripe committed ch_A but before the response returned to the
// application — a timeout — then ch_A is committed in Stripe
// and ch_B is committed when attempt 2 succeeds.
// Two charges, one customer, no Stripe-side error.
String idempotencyKey = UUID.randomUUID().toString(); // UUID_B on iteration 2
PaymentIntentCreateParams params = PaymentIntentCreateParams.builder()
.setAmount(amountCents)
.setCurrency("usd")
.setCustomer(customerId)
.setConfirm(true)
.build();
PaymentIntent pi = stripeClient.paymentIntents().create(
params,
RequestOptions.builder().setIdempotencyKey(idempotencyKey).build()
);
repo.save(new BillingRecord(
customerId, pi.getId(), billingPeriod, amountCents));
return pi;
} catch (StripeConnectException e) {
lastException = e;
attempt++;
if (attempt < 3) {
try {
// Cheap on virtual threads — no carrier thread blocked.
// This is genuinely fine from a threading perspective.
Thread.sleep(500L * (1L << attempt)); // 1s, 2s
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
throw new RuntimeException(ie);
}
}
} catch (StripeException e) {
throw new RuntimeException("Non-retryable Stripe exception", e);
}
}
throw new RuntimeException("Stripe exhausted after 3 attempts", lastException);
}, EXECUTOR);
}
}
The developer’s reasoning is internally coherent but wrong about the risk. Their argument: “If attempt 1 timed out and we send the same key on attempt 2, Stripe returns the ch_A result — the customer is charged once. If we send a new key, Stripe processes it as a new charge — ch_B. We don’t know if ch_A committed, so a new key is safer because it avoids acting on possibly-committed state.”
This reasoning inverts the risk. Stripe’s deduplication is specifically designed for the case where the outcome is unknown: a timeout on attempt 1 does not tell you whether Stripe committed ch_A. Sending the same key on attempt 2 is exactly the correct behavior: if ch_A committed, Stripe returns it without charging again; if ch_A did not commit, Stripe processes attempt 2 as a new charge. Sending a new key removes this safety net. The new-UUID approach guarantees a duplicate charge in the timeout case.
The correct framing of the timeout risk
A StripeConnectException with a timeout means the HTTP response was not received. Three things could have happened on Stripe’s side: (A) the request was never received by Stripe, (B) the request was received and is still processing, or (C) the request was received, processed, and Stripe sent a response that was lost in transit. In case (A), the same idempotency key on attempt 2 causes Stripe to process a fresh charge — correct. In case (C) — ch_A committed — the same key causes Stripe to return the ch_A result without processing again — correct. In case (B), Stripe will continue processing the original request and the second request with the same key will wait for the first to finish. In all three cases, sending the same key produces the correct outcome. Sending a new key produces a duplicate charge in case (C) — the most common timeout scenario for billing APIs operating at scale.
The fix: content-hash key computed outside the loop
// BillingWorker.java — fixed mode 2 (key computed outside loop)
@Component
public class BillingWorker {
private static final ExecutorService EXECUTOR =
Executors.newVirtualThreadPerTaskExecutor();
private final StripeClient stripeClient;
private final BillingRecordRepository repo;
public CompletableFuture<PaymentIntent> chargeAsync(
String customerId, long amountCents, String billingPeriod) {
// Key computed before the CompletableFuture lambda and the retry loop.
// Same customer + period + amount = same key across all attempts.
String idempotencyKey = contentHashKey(customerId, billingPeriod, amountCents);
return CompletableFuture.supplyAsync(() -> {
int attempt = 0;
StripeException lastException = null;
while (attempt < 3) {
try {
// idempotencyKey is captured from the enclosing scope.
// It is computed exactly once, before this loop.
// Every iteration — every virtual thread sleep — every retry —
// sends the same key. Stripe deduplicates correctly.
PaymentIntentCreateParams params = PaymentIntentCreateParams.builder()
.setAmount(amountCents)
.setCurrency("usd")
.setCustomer(customerId)
.setConfirm(true)
.build();
PaymentIntent pi = stripeClient.paymentIntents().create(
params,
RequestOptions.builder().setIdempotencyKey(idempotencyKey).build()
);
repo.save(new BillingRecord(
customerId, pi.getId(), billingPeriod, amountCents));
return pi;
} catch (StripeConnectException e) {
lastException = e;
attempt++;
if (attempt < 3) {
try {
Thread.sleep(500L * (1L << attempt));
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
throw new RuntimeException(ie);
}
}
} catch (StripeException e) {
throw new RuntimeException("Non-retryable Stripe exception", e);
}
}
throw new RuntimeException("Stripe exhausted after 3 attempts", lastException);
}, EXECUTOR);
}
private String contentHashKey(String customerId, String period, long amountCents) {
String input = customerId + "|" + period + "|" + amountCents;
try {
MessageDigest md = MessageDigest.getInstance("SHA-256");
byte[] hash = md.digest(input.getBytes(StandardCharsets.UTF_8));
StringBuilder sb = new StringBuilder();
for (byte b : hash) sb.append(String.format("%02x", b));
return sb.toString();
} catch (NoSuchAlgorithmException e) {
throw new IllegalStateException("SHA-256 not available", e);
}
}
}
The content-hash key also solves a second failure mode that is specific to virtual-thread batch billing: if chargeAsync() is called concurrently for the same customer from two different code paths (e.g., a scheduled job and an admin retry endpoint), both calls compute the same key and Stripe deduplicates correctly. This would not be true if keys were random UUIDs.
Migration from reactive retry operators
If migrating from a Spring WebClient reactive retry implementation, the content-hash key pattern from the reactive version carries over directly. The key in reactive code should have been computed outside the retryWhen() operator (outside the reactive publisher that gets re-subscribed on retry). The same principle applies to virtual-thread code: compute outside the retry loop.
// BEFORE: reactive (Spring WebClient)
public Mono<PaymentIntent> chargeReactive(String customerId, long amountCents, String period) {
String idempotencyKey = contentHashKey(customerId, period, amountCents); // outside retryWhen
return webClient.post()
.uri("/v1/payment_intents")
.header("Idempotency-Key", idempotencyKey)
.bodyValue(buildParams(customerId, amountCents))
.retrieve()
.bodyToMono(PaymentIntent.class)
.retryWhen(Retry.backoff(3, Duration.ofMillis(500))
.filter(e -> e instanceof WebClientRequestException));
}
// AFTER: virtual threads (blocking)
public CompletableFuture<PaymentIntent> chargeVirtualThread(
String customerId, long amountCents, String period) {
String idempotencyKey = contentHashKey(customerId, period, amountCents); // outside loop
return CompletableFuture.supplyAsync(() -> {
// ... retry loop with Thread.sleep() ...
// idempotencyKey captured from enclosing scope — correct in both models
}, EXECUTOR);
}
Mode 3: StructuredTaskScope.ShutdownOnFailure + outer retry loop — scope shutdown cancels all running forks — outer retry restarts all forks with new UUIDs — ch_B per already-charged customer
Java 21 introduced structured concurrency as a preview API (java.util.concurrent.StructuredTaskScope, finalized in Java 23). Structured concurrency provides a disciplined model for forking parallel tasks: a StructuredTaskScope represents a lifetime boundary — all forked tasks must complete before the scope exits. ShutdownOnFailure is a built-in policy that cancels all running forks when any fork throws an exception, then throws the exception to the scope owner when scope.join().throwIfFailed() is called.
Spring Boot applications on Java 21+ can use StructuredTaskScope directly for parallel billing. The pattern is appealing for batch billing: fork one task per customer, wait for all, collect results. If any task fails, ShutdownOnFailure cancels all running tasks and surfaces the exception. The outer retry loop creates a new scope and re-forks all tasks.
// BatchBillingService.java — unsafe mode 3 (StructuredTaskScope + outer retry loop)
@Service
public class BatchBillingService {
private final StripeClient stripeClient;
private final BillingRecordRepository repo;
public void billAllCustomers(List<CustomerBillingItem> customers, String billingPeriod)
throws InterruptedException {
int attempt = 0;
Exception lastException = null;
while (attempt < 3) {
try (var scope = new StructuredTaskScope.ShutdownOnFailure()) {
List<StructuredTaskScope.Subtask<PaymentIntent>> subtasks = new ArrayList<>();
for (CustomerBillingItem customer : customers) {
subtasks.add(scope.fork(() -> {
// Developer's reasoning:
// "StructuredTaskScope guarantees that all forks complete
// before the scope exits — ShutdownOnFailure shuts down
// the scope if any fork fails. On retry, a new scope is
// created and new forks are submitted. UUID is inside the fork
// lambda — each customer gets a fresh key. This is correct
// because on the outer retry, we don't know which customers
// succeeded, and generating a new key for already-committed
// charges is safe — Stripe returns the same result if the
// charge went through."
//
// The problem: Stripe deduplication is by key, not by customer.
// A new UUID on the outer retry is a NEW charge request, not a
// lookup. If customer A's fork committed ch_A before customer B's
// fork failed (triggering ShutdownOnFailure), the outer retry
// creates a new scope — customer A's new fork sends UUID_B —
// Stripe processes ch_B — two charges for customer A.
// ShutdownOnFailure cancels RUNNING forks (those not yet finished).
// Forks that COMPLETED successfully have already committed to Stripe.
// The scope cannot un-commit them.
String idempotencyKey = UUID.randomUUID().toString(); // UUID_B on outer retry
PaymentIntentCreateParams params = PaymentIntentCreateParams.builder()
.setAmount(customer.getAmountCents())
.setCurrency("usd")
.setCustomer(customer.getStripeCustomerId())
.setConfirm(true)
.build();
PaymentIntent pi = stripeClient.paymentIntents().create(
params,
RequestOptions.builder().setIdempotencyKey(idempotencyKey).build()
);
repo.save(new BillingRecord(
customer.getCustomerId(), pi.getId(),
billingPeriod, customer.getAmountCents()));
return pi;
}));
}
scope.join().throwIfFailed();
return; // all forks succeeded
} catch (ExecutionException | StripeException e) {
lastException = e;
attempt++;
if (attempt < 3) {
try {
Thread.sleep(1000L * attempt);
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
throw ie;
}
}
}
}
throw new RuntimeException("Batch billing failed after 3 attempts", lastException);
}
}
The developer’s reasoning contains a plausible but incorrect claim about Stripe deduplication. Stripe does not deduplicate by customer+amount+period. Stripe deduplicates by idempotency key. If customer A is charged with key UUID_A (ch_A committed) and then charged again with key UUID_B in the outer retry loop, Stripe has no way to know that UUID_B is a retry of UUID_A — it processes a new charge. ShutdownOnFailure’s scope shutdown cancels forks that are still running; it has no ability to roll back or cancel forks that already completed, because those forks already made HTTP calls to Stripe that returned successfully.
Anatomy of the failure window
The failure window opens when customer B’s fork fails after customer A’s fork has already committed. The timeline:
- Attempt 1, scope created. All 100 customer forks submitted. Virtual threads schedule them concurrently.
- Customers 1–80 complete successfully. ch_A committed in Stripe. DB records written.
- Customer 81’s fork throws
StripeConnectException.ShutdownOnFailurecancels the 19 running forks (customers 82–100 in various states of execution). The scope exits. - Outer retry: new scope. 100 customers again. Customers 1–80 forks send UUID_B. Stripe commits ch_B for each. DB records written (or fail on unique constraint if the DB has one).
- Customer 81 retried with UUID_B — correct (ch_A was not committed). Customer 82–100 retried — may or may not have committed in the cancelled forks.
The blast radius is proportional to how many forks completed successfully before the failure. In a 100-customer batch where 80 completed before the failure, the outer retry creates 80 duplicate charges. The charges are not detectable by Spring or the Stripe SDK — both ch_A and ch_B return HTTP 200 from Stripe.
Why ShutdownOnFailure does not protect against this
ShutdownOnFailure is designed to cancel future work, not to undo completed work. A fork that finishes successfully exits the scope’s tracking. From the scope’s perspective, the fork is done — there is nothing to cancel. ShutdownOnFailure cancels forks that are still executing (pending I/O, sleeping, or in their task body) by interrupting their virtual threads. Forks that returned a result are unaffected.
This is not a bug in StructuredTaskScope. Structured concurrency does not provide transactional semantics over external side effects (HTTP calls, database writes). It provides a scope boundary for task lifetimes. Idempotency for external side effects is the application’s responsibility, not the scope’s.
Fix option A: content-hash key per fork, computed outside the retry loop
Computing the key outside the outer retry loop ensures the same key is submitted on every attempt. Stripe deduplicates correctly for forks whose previous attempt committed ch_A.
// BatchBillingService.java — fixed mode 3, option A (keys computed outside retry loop)
@Service
public class BatchBillingService {
private final StripeClient stripeClient;
private final BillingRecordRepository repo;
public void billAllCustomers(List<CustomerBillingItem> customers, String billingPeriod)
throws InterruptedException {
// Keys computed once, outside the retry loop.
// Same customer + period + amount = same key on every attempt.
// Stripe deduplicates based on these stable keys.
Map<String, String> idempotencyKeys = customers.stream()
.collect(Collectors.toMap(
CustomerBillingItem::getCustomerId,
c -> contentHashKey(c.getCustomerId(), billingPeriod, c.getAmountCents())
));
int attempt = 0;
Exception lastException = null;
while (attempt < 3) {
try (var scope = new StructuredTaskScope.ShutdownOnFailure()) {
for (CustomerBillingItem customer : customers) {
// Key from the pre-computed map — stable across all outer retry attempts.
String idempotencyKey = idempotencyKeys.get(customer.getCustomerId());
scope.fork(() -> {
PaymentIntentCreateParams params = PaymentIntentCreateParams.builder()
.setAmount(customer.getAmountCents())
.setCurrency("usd")
.setCustomer(customer.getStripeCustomerId())
.setConfirm(true)
.build();
PaymentIntent pi = stripeClient.paymentIntents().create(
params,
RequestOptions.builder().setIdempotencyKey(idempotencyKey).build()
);
repo.save(new BillingRecord(
customer.getCustomerId(), pi.getId(),
billingPeriod, customer.getAmountCents()));
return pi;
});
}
scope.join().throwIfFailed();
return;
} catch (ExecutionException | StripeException e) {
lastException = e;
attempt++;
if (attempt < 3) {
try {
Thread.sleep(1000L * attempt);
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
throw ie;
}
}
}
}
throw new RuntimeException("Batch billing failed after 3 attempts", lastException);
}
}
Fix option B: skip already-committed customers on retry
A complementary fix is to query the database before each retry attempt and skip customers who already have a committed BillingRecord for the billing period. This reduces the number of Stripe calls on retry to only the customers who were not yet charged.
// BatchBillingService.java — fixed mode 3, option B (skip committed on retry)
// Combined with option A for belt-and-suspenders defense
private Set<String> alreadyCommitted(List<CustomerBillingItem> customers, String period) {
// Query DB for customers with a committed BillingRecord for this period.
return customers.stream()
.map(CustomerBillingItem::getCustomerId)
.filter(id -> repo.existsByCustomerIdAndBillingPeriod(id, period))
.collect(Collectors.toSet());
}
// In the retry loop:
// Set<String> committed = alreadyCommitted(customers, billingPeriod);
// List<CustomerBillingItem> remaining = customers.stream()
// .filter(c -> !committed.contains(c.getCustomerId()))
// .toList();
// ... fork only remaining customers ...
The DB skip reduces cost but does not eliminate the duplicate charge risk by itself — there is a race window between the DB query and the Stripe call. Option B alone is insufficient; the combination of stable content-hash keys (option A) plus DB skip (option B) provides the strongest defense: Stripe deduplicates if the DB is out of sync, and the DB skip reduces unnecessary Stripe calls.
Test: verifying keys are stable across StructuredTaskScope outer retry
// BatchBillingServiceTest.java — JUnit 5 + WireMock
@SpringBootTest
class BatchBillingServiceTest {
@Autowired
BatchBillingService batchBillingService;
@RegisterExtension
static WireMockExtension wireMock = WireMockExtension.newInstance()
.options(wireMockConfig().dynamicPort())
.build();
@Test
void outerRetryUsesTheSameKeyForAlreadyChargedCustomers() throws Exception {
// customer_A succeeds on attempt 1.
// customer_B fails on attempt 1, triggers scope shutdown.
// customer_A re-submitted on attempt 2 with the same key.
// Assert: the Idempotency-Key for customer_A is identical on both attempts.
AtomicInteger customerACallCount = new AtomicInteger(0);
wireMock.stubFor(post(urlPathEqualTo("/v1/payment_intents"))
.withRequestBody(matchingJsonPath("$.customer", equalTo("cus_A")))
.willReturn(aResponse().withTransformer(new ResponseDefinitionTransformer() {
@Override
public ResponseDefinition transform(Request request,
ResponseDefinition responseDefinition, FileSource files,
Parameters parameters) {
customerACallCount.incrementAndGet();
return ok().withBody(paymentIntentJson("pi_A")).build();
}
@Override public String getName() { return "customer-a-counter"; }
})));
// customer_B always fails
wireMock.stubFor(post(urlPathEqualTo("/v1/payment_intents"))
.withRequestBody(matchingJsonPath("$.customer", equalTo("cus_B")))
.inScenario("cus_b_retry").whenScenarioStateIs(STARTED)
.willReturn(serverError()).willSetStateTo("ok"));
wireMock.stubFor(post(urlPathEqualTo("/v1/payment_intents"))
.withRequestBody(matchingJsonPath("$.customer", equalTo("cus_B")))
.inScenario("cus_b_retry").whenScenarioStateIs("ok")
.willReturn(ok().withBody(paymentIntentJson("pi_B"))));
List<CustomerBillingItem> customers = List.of(
new CustomerBillingItem("cus_A", "cus_stripe_A", 5000L),
new CustomerBillingItem("cus_B", "cus_stripe_B", 7500L)
);
batchBillingService.billAllCustomers(customers, "2026-10");
// Collect all Idempotency-Key headers for customer_A
List<String> keysForA = wireMock
.findAll(postRequestedFor(urlPathEqualTo("/v1/payment_intents"))
.withRequestBody(matchingJsonPath("$.customer", equalTo("cus_A"))))
.stream()
.map(r -> r.getHeader("Idempotency-Key"))
.collect(Collectors.toList());
// customer_A was called on both attempts
assertThat(keysForA).hasSize(2);
// Both attempts used the same key — Stripe deduplication active
assertThat(keysForA.get(0)).isEqualTo(keysForA.get(1));
}
}
Comparison: failure modes across all three virtual threads + Stripe patterns
| Mode | Retry mechanism | Re-execution unit | UUID position (unsafe) | Developer misconception |
|---|---|---|---|---|
| 1 | @Retryable AOP, virtual-thread executor |
Entire @Retryable method body (fresh invocation via proceed()) |
Inside the @Retryable method |
Virtual threads use continuations like coroutines — retry might resume preserved state |
| 2 | Manual while loop + Thread.sleep() |
Entire loop body per iteration | Inside the while loop body |
New UUID on retry is safer because we don’t know if ch_A committed |
| 3 | StructuredTaskScope outer retry loop |
New scope + all fork lambdas per outer retry attempt | Inside the fork lambda (inside the scope) | ShutdownOnFailure manages task lifecycles — it won’t re-run already-committed forks |
| Mode | Stripe impact | DB impact | Detection signal |
|---|---|---|---|
| 1 | ch_B per customer on retry; both ch_A and ch_B return HTTP 200 | Duplicate BillingRecord rows if no unique constraint on (customerId, period) |
Stripe dashboard: two PaymentIntents same customer same period. Application log: @Retryable retry event. |
| 2 | ch_B on iteration 2 if ch_A committed on iteration 1 timeout; no signal in Stripe API response | Duplicate row or unique constraint violation on second loop iteration | Stripe: two charges same amount same customer. Application: no error (ch_B succeeds). DB: duplicate row or exception on insert. |
| 3 | ch_B per already-charged customer; blast radius = number of successful forks before failing fork | DB records written for both attempts; duplicate rows if no unique constraint | Stripe: N duplicate charges for N successful forks. Application: no error. DB: N duplicate rows or unique constraint violations. |
Virtual threads vs. Kotlin coroutines: where the idempotency behavior actually differs
Developers migrating from Kotlin coroutines to Java virtual threads, or comparing the two, often ask whether the idempotency behavior is identical. The answer is: the failure modes in this post have structural parallels in Kotlin coroutines, but the mechanisms are different.
In Kotlin, Flow.retry{} re-collects the upstream flow{} builder — re-executing the builder block including any UUID.randomUUID() calls inside it. This is structurally identical to Mode 1 above: a retry mechanism re-executes a code block that contains UUID generation. The fix is the same: compute the key outside the retried block. See the Kotlin Flow + Exposed + Stripe post for the coroutine-specific variants.
The meaningful difference between virtual threads and coroutines for Stripe idempotency is not in the retry semantics (both re-execute) but in the cancellation behavior. A cancelled Kotlin coroutine (CancellationException) at a suspension point does not necessarily mean the Stripe HTTP request was not committed — the request may already have been sent. Virtual thread interruption via StructuredTaskScope.ShutdownOnFailure has the same property: a cancelled virtual thread does not roll back an HTTP request already in flight. In both models, the Stripe charge is a committed external side effect that is independent of the caller’s execution lifecycle.
Spring Boot 3.2 virtual threads migration checklist for Stripe integration
When enabling spring.threads.virtual.enabled=true on an existing Spring Boot application that uses Stripe, audit the following:
@Retryablemethods: verify thatUUID.randomUUID()is not inside any@Retryablemethod body. Use the content-hash key pattern or pass the key as a parameter.- Manual retry loops: any code that was migrated from reactive retry operators (
Mono.retryWhen,Observable.retry) to awhileloop withThread.sleep()— check that UUID generation is outside the loop. @Asyncmethods: with virtual threads enabled,@Asyncmethods no longer require a thread pool — they run on virtual threads. If the@Asyncmethod previously relied on a fixed thread pool executor to limit Stripe concurrency, that limit is now gone. A@Scheduledtask that calls@Asyncfor each customer can now launch as many virtual thread tasks as customers. Add an explicit semaphore or rate-limiter if needed.- Idempotency key generation location: regardless of thread model, the universal rule is: compute the key before the boundary where re-execution can occur. That boundary is the
@RetryableAOP proxy, the retry loop boundary, or the fork submission in aStructuredTaskScope. - Database unique constraints: add a unique constraint on
(customer_id, billing_period)to thebilling_recordstable as a last-resort duplicate charge detector. A duplicateINSERTwill throw an exception instead of silently creating a second record.
Why Keybrake helps here
Keybrake sits between your application and the Stripe API, enforcing per-agent and per-run spend caps via a proxy. If a virtual-thread retry loop generates duplicate charges, Keybrake’s per-day spend cap for the issuing agent key will fire before the N-th duplicate clears. The audit log records every proxied request — multiple requests with different Idempotency-Key values for the same customer in the same minute are a visible anomaly in the Keybrake dashboard, even when they produce valid Stripe responses. This does not replace content-hash keys, but it provides an operational early-warning layer that content-hash keys alone cannot.
The detection gap the Stripe dashboard has: two PaymentIntents created in quick succession for the same customer appear as two separate line items. Without the idempotency key linkage, there is no Stripe-native signal that one is a retry of the other. Keybrake sees the sequence of requests keyed to a specific agent key and billing period, making the anomaly pattern detectable at the proxy layer before it becomes a customer support problem.
See also: Spring Batch + Stripe idempotency, Spring @Async + @Transactional + Stripe, Kotlin Coroutines @Transactional + Stripe, and Spring WebClient reactive retry + Stripe.
Put a spend cap on your agent’s Stripe key
Keybrake proxies your Stripe calls with per-agent spend caps, endpoint allowlists, and a full audit log. One-click revoke stops a runaway retry loop before it clears your daily cap.