Spring Boot, Spring Batch, and Stripe Integration: How FaultTolerantStep Retry Re-executes ItemWriter.write(), @Scheduled Re-launch Creates a New JobInstance That Charges All Items Again, and Skip-Scan Logic Re-processes Items Through ItemProcessor and Regenerates Idempotency Keys
Spring Batch’s fault-tolerance infrastructure — chunk retry, skip, and job restart — is designed to make batch jobs robust against transient failures. Applied to Stripe charges, these same mechanisms generate duplicate charges in ways that are invisible to both the application log and the Stripe dashboard, because each failure-recovery path re-invokes code that should have been idempotent but was not.
This post covers three Spring Batch + Stripe integration failure modes that are structurally distinct from the basic Spring Boot @Transactional post, the Spring WebClient reactive retry post, and the Spring @Async + @Transactional post. The modes here are specific to Spring Batch’s chunk-oriented processing model, its FaultTolerantStep retry and skip mechanics, and the relationship between JobInstance, JobExecution, and the @Scheduled trigger pattern that most teams use to schedule batch jobs.
Background: Spring Batch chunk processing, fault tolerance, and the Stripe idempotency contract
Spring Batch’s chunk-oriented processing model reads items from an ItemReader, passes them through an ItemProcessor, and writes them in batches via an ItemWriter. The commit-interval controls how many items are accumulated before write() is called. The entire read-process-write cycle for a chunk runs inside a single Spring transaction: if write() throws an exception, the transaction rolls back, and no DB writes from that chunk are committed.
Spring Batch’s FaultTolerantStep adds two capabilities on top of this model: retry and skip. Retry: when write() throws a retryable exception, the step re-invokes write() with the same chunk, up to the configured retry limit. Skip: when retries are exhausted, the step switches to single-item processing to identify which item caused the exception and mark it as skipped. Each of these mechanisms re-invokes ItemWriter.write() — and in the case of skip-scan, also re-invokes ItemProcessor.process().
Spring Batch’s JobInstance and JobExecution are two distinct concepts. A JobInstance is identified by the job name plus JobParameters. A JobExecution is one attempt to run a JobInstance. A failed JobExecution can be restarted via JobOperator.restart(), which creates a new JobExecution against the same JobInstance and resumes from the last committed chunk. A new JobInstance — created when jobLauncher.run(job, differentParameters) is called — starts from the beginning, with no knowledge of prior executions.
Stripe’s idempotency key contract: Stripe deduplicates requests based on the Idempotency-Key header. Two requests with the same key for the same endpoint within 24 hours return the cached result of the first successful request. Two requests with different keys for the same customer at the same amount on the same day are two charges. The contract is per-key, per-endpoint, not per-customer, not per-amount.
Mode 1: FaultTolerantStep retry calls ItemWriter.write() again — UUID inside write() generates UUID_B on the retry invocation — ch_B alongside committed ch_A
When a FaultTolerantStep is configured with .retry(StripeConnectException.class).retryLimit(3), Spring Batch wraps the ItemWriter.write() call in a RetryTemplate. If write() throws a retryable exception, RetryTemplate calls the retry callback again — which calls write() again with the same list of items. This is a new method invocation: the JVM creates a new stack frame for write(), and every local variable inside the method is initialized from scratch.
// BillingItemWriter.java — unsafe mode 1 (UUID inside write())
@Component
public class BillingItemWriter implements ItemWriter<CustomerBillingItem> {
private final StripeClient stripeClient;
private final BillingRecordRepository repo;
@Override
public void write(List<? extends CustomerBillingItem> items) throws Exception {
for (CustomerBillingItem item : items) {
// Developer's reasoning:
// "Each item gets a fresh UUID. This is correct — we don't want
// two items sharing an idempotency key. The UUID is local to the
// loop iteration. This is the standard pattern."
//
// The problem: write() is called again by RetryTemplate on
// StripeConnectException. Each retry invocation of write() creates
// a new stack frame. Every local variable re-initializes.
// The UUID for every item in the chunk regenerates on every retry call.
// Attempt 1: UUID_A sent to Stripe. Stripe processes ch_A.
// Network timeout before response returns. StripeConnectException thrown.
// RetryTemplate catches it. Attempt 2: write() invoked again.
// UUID_B generated. Stripe sees a new key — processes ch_B.
// Two charges for the same customer, same billing period.
String idempotencyKey = UUID.randomUUID().toString(); // UUID_B on retry
PaymentIntentCreateParams params = PaymentIntentCreateParams.builder()
.setAmount(item.getAmountCents())
.setCurrency("usd")
.setCustomer(item.getStripeCustomerId())
.setConfirm(true)
.build();
PaymentIntent pi = PaymentIntent.create(params,
RequestOptions.builder().setIdempotencyKey(idempotencyKey).build());
repo.save(new BillingRecord(item.getCustomerId(), pi.getId(),
item.getBillingPeriod(), item.getAmountCents()));
}
}
}
// BillingJobConfig.java — FaultTolerantStep with retry
@Bean
public Step billingStep(StepBuilderFactory stepBuilderFactory,
ItemReader<CustomerBillingItem> reader,
ItemProcessor<CustomerBillingItem, CustomerBillingItem> processor,
BillingItemWriter writer) {
return stepBuilderFactory.get("billingStep")
.<CustomerBillingItem, CustomerBillingItem>chunk(50)
.reader(reader)
.processor(processor)
.writer(writer)
.faultTolerant()
.retry(StripeConnectException.class)
.retryLimit(3)
.build();
}
The developer’s mental model has one accurate component and one fatal flaw. Accurate: UUID is local to the loop iteration, so two items in the same chunk do not share a key. Fatal flaw: write() itself is re-invoked on retry, not just re-attempted by the JVM for the same invocation. The retryLimit of 3 means up to three distinct invocations of write() for the same chunk, each with fresh local variables, each producing new UUIDs for every item in the chunk.
Why the retry fires specifically on Stripe + DB patterns
The failure window is open whenever Stripe’s HTTP response is delayed enough to exceed the application’s HTTP client timeout, or whenever the Stripe API returns a 500 or 503 that the Stripe Java SDK maps to StripeConnectException. In a billing job that processes hundreds or thousands of customers, the probability of at least one Stripe call hitting a transient error within the job run is higher than in a single request-response context. The retry is triggered per chunk, not per item: all 50 items in the chunk get new UUIDs on retry, not just the one that caused the exception.
There is an additional subtlety: Stripe’s own SDK performs internal retries for network errors before surfacing the exception to the application. By default, the Stripe Java SDK retries up to two times on StripeConnectException. If the first SDK retry succeeds (Stripe committed the charge), the SDK returns the charge to the application without throwing. If all SDK retries fail, the SDK surfaces the exception. At the point Spring Batch’s RetryTemplate receives the exception, the Stripe charge may or may not be committed — the SDK’s internal retry behavior is opaque to Spring Batch. This means the Spring Batch retry may call write() with UUID_B even when ch_A was committed during the SDK’s internal retry on the first write() invocation.
The fix: content-hash key computed before entering write()
Moving the idempotency key computation out of write() — into the ItemProcessor or into the CustomerBillingItem object itself — ensures the key is stable across all retry invocations of write(). The processor runs once per item per chunk, before the retry loop. Keys computed there are captured in the CustomerBillingItem output object and passed unchanged to every retry invocation of write().
// BillingItemProcessor.java — key computed in processor (stable across write() retries)
@Component
public class BillingItemProcessor
implements ItemProcessor<CustomerBillingItem, CustomerBillingItem> {
@Override
public CustomerBillingItem process(CustomerBillingItem item) throws Exception {
// Content-hash key computed here — in the processor, before the retry loop.
// RetryTemplate wraps write() — not process(). This runs exactly once per item
// per chunk, regardless of how many times write() is retried.
// The key is embedded in the output object and reused by every write() invocation.
String idempotencyKey = contentHashKey(
item.getCustomerId(),
item.getBillingPeriod(),
item.getAmountCents()
);
item.setIdempotencyKey(idempotencyKey);
return item;
}
private String contentHashKey(String customerId, String period, long amountCents) {
String input = customerId + "|" + period + "|" + amountCents;
try {
MessageDigest md = MessageDigest.getInstance("SHA-256");
byte[] hash = md.digest(input.getBytes(StandardCharsets.UTF_8));
StringBuilder sb = new StringBuilder();
for (byte b : hash) sb.append(String.format("%02x", b));
return sb.toString();
} catch (NoSuchAlgorithmException e) {
throw new IllegalStateException("SHA-256 not available", e);
}
}
}
// BillingItemWriter.java — fixed mode 1 (key from processor output)
@Component
public class BillingItemWriter implements ItemWriter<CustomerBillingItem> {
private final StripeClient stripeClient;
private final BillingRecordRepository repo;
@Override
public void write(List<? extends CustomerBillingItem> items) throws Exception {
for (CustomerBillingItem item : items) {
// Key is from the processor output — computed before the retry loop.
// Attempt 1: hash("cus_A|2026-10|5000") = "a3f7..." → Stripe commits ch_A.
// Network timeout. RetryTemplate retries write().
// Attempt 2: same key "a3f7..." — Stripe deduplicates — returns ch_A result.
// No ch_B. Customer charged exactly once.
String idempotencyKey = item.getIdempotencyKey(); // stable across retries
PaymentIntentCreateParams params = PaymentIntentCreateParams.builder()
.setAmount(item.getAmountCents())
.setCurrency("usd")
.setCustomer(item.getStripeCustomerId())
.setConfirm(true)
.build();
PaymentIntent pi = PaymentIntent.create(params,
RequestOptions.builder().setIdempotencyKey(idempotencyKey).build());
repo.save(new BillingRecord(item.getCustomerId(), pi.getId(),
item.getBillingPeriod(), item.getAmountCents()));
}
}
}
The content-hash key approach has an additional benefit for Spring Batch specifically: if the same customer appears in the reader multiple times across different job runs (due to a cursor reset, a step restart, or a reader bug), the same key is generated every time. Stripe deduplicates not just within a single job run but across all runs within the 24-hour deduplication window. This provides a belt-and-suspenders guard against the mode 2 failure described below.
Disabling the Stripe SDK’s internal retry to clarify ownership
When both the Stripe SDK and Spring Batch perform retries, two independent retry loops operate on the same Stripe call. The Stripe SDK’s internal retry loop (default: 2 attempts) runs first; if it succeeds, the exception is swallowed and Spring Batch’s retry never fires. If the SDK exhausts its retries, Spring Batch fires its retry with a new write() invocation. Disabling the Stripe SDK’s internal retry — via StripeClient.builder().setMaxNetworkRetries(0) — and letting Spring Batch’s RetryTemplate own all retry decisions simplifies the retry topology and makes the behavior auditable from a single configuration point. Combined with content-hash keys, this is the recommended configuration for Spring Batch billing jobs.
// StripeConfig.java — disable SDK internal retry; let Spring Batch own all retries
@Configuration
public class StripeConfig {
@Bean
public StripeClient stripeClient(@Value("${stripe.api.key}") String apiKey) {
return StripeClient.builder()
.setApiKey(apiKey)
.setMaxNetworkRetries(0) // Spring Batch RetryTemplate owns all retries
.build();
}
}
Mode 2: @Scheduled + jobLauncher.run(job, new JobParametersBuilder().addLong(“run.id”, currentTimeMillis())) creates a new JobInstance instead of restarting the failed execution — all items re-processed from scratch — customers already charged in the failed run receive ch_B
Spring Batch prevents a JobInstance from being run more than once successfully. If you call jobLauncher.run(job, params) with the same JobParameters for a job that already completed successfully, Spring Batch throws JobInstanceAlreadyCompleteException. For jobs triggered on a schedule — monthly billing, weekly reconciliation, nightly exports — this behavior causes a problem: the job succeeds on the first of the month, and the next @Scheduled invocation on the second of the month fails with JobInstanceAlreadyCompleteException because the parameters haven’t changed.
The standard workaround documented in the Spring Batch reference, Stack Overflow, and virtually every tutorial is to add a unique parameter to the JobParameters — typically a timestamp: new JobParametersBuilder().addLong(“run.id”, System.currentTimeMillis()).toJobParameters(). This creates a new JobInstance on each scheduled invocation, bypassing the already-complete check.
The problem with this pattern becomes apparent when a job fails mid-run. The failed run has processed and charged some customers. The next @Scheduled invocation creates a brand new JobInstance — because run.id is a different timestamp. Spring Batch has no concept of “restart this failed execution” for a new JobInstance. It starts from the beginning of the step. Every customer in the reader’s data set is processed again.
// BillingScheduler.java — unsafe mode 2 (@Scheduled + new JobInstance per run)
@Component
public class BillingScheduler {
private final JobLauncher jobLauncher;
private final Job billingJob;
@Scheduled(cron = "0 0 2 1 * *") // 2 AM on the 1st of every month
public void runBillingJob() throws JobExecutionException {
// Developer's reasoning:
// "Adding run.id is the standard pattern for allowing @Scheduled to
// run the same job repeatedly. The timestamp ensures we get a new
// JobInstance each month. If the job fails and @Scheduled fires again,
// the new run picks up where the old one left off — that's Spring Batch's
// restart semantics."
//
// The problem: restart semantics apply to a FAILED JobExecution being
// re-run against the SAME JobInstance with the SAME JobParameters.
// A new run.id = a new JobInstance = no restart.
// Spring Batch has no memory of the failed execution for this new JobInstance.
// The step starts at chunk 0, item 0, regardless of what the failed execution
// processed. Every customer in the reader is processed again.
// For a billing job with 10,000 customers where 8,000 were charged before failure:
// the new JobInstance charges all 10,000 — 8,000 of whom now have ch_A and ch_B.
JobParameters params = new JobParametersBuilder()
.addLong("run.id", System.currentTimeMillis()) // new JobInstance every run
.addString("billingPeriod", YearMonth.now().toString())
.toJobParameters();
jobLauncher.run(billingJob, params);
}
}
The misconception here is subtle because it conflates two separate Spring Batch concepts: the ability to run the same job definition more than once (which the run.id trick enables by creating new JobInstances), and the ability to restart a failed execution from its last commit point (which requires the same JobParameters, hence the same JobInstance). The run.id trick solves the first problem by working around the second.
Why this fires specifically for batch billing
Batch billing jobs are the worst-case scenario for this failure mode because the reader is typically unbounded from Spring Batch’s perspective. A JdbcCursorItemReader that reads all active subscribers from the database will re-read all 10,000 customers on the new JobInstance, including the 8,000 who were successfully charged in the failed execution. If the ItemWriter uses UUID.randomUUID() inside write() (Mode 1) or inside ItemProcessor.process(), 8,000 customers receive a second charge with UUID_B alongside the UUID_A charge from the first execution.
The failure is exacerbated by timing: the @Scheduled cron may fire again within minutes of the failed execution (if the retry interval is short), within hours (if nightly), or at the next scheduled period (if monthly). In all cases, Spring Batch proceeds without warning because a new JobInstance is a valid, fresh execution from Spring Batch’s perspective — it has never run before.
Fix option A: use JobOperator.restart() to resume the failed execution
The correct restart mechanism for a failed execution is JobOperator.restart(failedExecutionId). This resumes the failed JobExecution from its last committed chunk, against the same JobInstance. Customers whose chunks committed successfully in the failed execution are not re-read by the reader (Spring Batch advances the reader cursor past committed chunks on restart).
// BillingScheduler.java — fix option A: restart failed execution, launch new if none exists
@Component
public class BillingScheduler {
private final JobLauncher jobLauncher;
private final JobOperator jobOperator;
private final JobExplorer jobExplorer;
private final Job billingJob;
@Scheduled(cron = "0 0 2 1 * *")
public void runBillingJob() throws JobExecutionException {
String billingPeriod = YearMonth.now().toString();
// Check if there is a failed execution for this billing period's JobInstance.
// If so, restart it — resume from the last committed chunk.
// If not, launch a fresh execution.
JobParameters params = new JobParametersBuilder()
.addString("billingPeriod", billingPeriod)
// No run.id — same parameters = same JobInstance = restart semantics apply
.toJobParameters();
List<JobInstance> instances = jobExplorer.getJobInstances(
billingJob.getName(), 0, 1);
if (!instances.isEmpty()) {
List<JobExecution> executions = jobExplorer.getJobExecutions(instances.get(0));
for (JobExecution execution : executions) {
if (execution.getStatus() == BatchStatus.FAILED) {
// Restart the failed execution — resumes from last committed chunk.
// Customers in committed chunks are not re-read or re-charged.
jobOperator.restart(execution.getId());
return;
}
}
}
// No failed execution — fresh launch for this period.
jobLauncher.run(billingJob, params);
}
}
The trade-off: without run.id, Spring Batch will throw JobInstanceAlreadyCompleteException if you call jobLauncher.run() for a JobInstance that has already completed successfully. This is the desired behavior for billing: once all customers are charged for a given period, re-running the job for the same period should be a no-op or an error, not a re-charge. The fix makes this explicit rather than papering over it with a timestamp parameter.
Fix option B: content-hash key as a catch-all deduplication layer
In systems where changing the scheduler is impractical in the short term, content-hash keys provide a catch-all deduplication layer at the Stripe layer. Even if the new JobInstance re-processes all 10,000 customers, Stripe deduplicates charges whose key is unchanged within the 24-hour window. For monthly billing, all charges from the first execution and all charges from the re-run share the same content-hash key (same customerId, same billingPeriod, same amountCents). Stripe returns the cached result of the first successful charge for the already-charged customers. The content-hash key fix is complementary to, not a replacement for, the scheduler fix: it prevents duplicate charges regardless of how many times the job runs.
// Content-hash key for monthly billing — stable across all executions for the same period
private String billingKey(String customerId, String billingPeriod, long amountCents) {
// Key is identical for cus_A|2026-10|5000 regardless of:
// — Which JobInstance computed it
// — Which JobExecution computed it
// — How many times write() was called
// — Which node in a clustered batch environment processed the item
// Stripe deduplicates all requests with this key within 24 hours.
String input = customerId + "|" + billingPeriod + "|" + amountCents;
try {
byte[] hash = MessageDigest.getInstance("SHA-256")
.digest(input.getBytes(StandardCharsets.UTF_8));
StringBuilder sb = new StringBuilder();
for (byte b : hash) sb.append(String.format("%02x", b));
return sb.toString();
} catch (NoSuchAlgorithmException e) {
throw new IllegalStateException(e);
}
}
The multi-period billing hazard: billing period in the key
A billing key that omits the billing period — using only customerId + "|" + amountCents — deduplicates across periods as well as within a period. A customer charged $50 in September and $50 in October would have the same key in both months, and Stripe would return the September charge result for the October charge request. The billing period must always be part of the content-hash inputs. If the amount can change between periods (subscription upgrade mid-period), the amount must be included too. The key must encode every dimension that distinguishes two logically different charges.
Mode 3: FaultTolerantStep skip-scan logic re-processes items through ItemProcessor — UUID computed in process() generates UUID_B — write() sends ch_B alongside committed ch_A
Spring Batch’s FaultTolerantStep handles the case where a chunk-level exception cannot be resolved by retry. When write() throws a non-retryable exception, or when the retry limit is exhausted, Spring Batch enters “skip-scan” mode: it processes each item in the failed chunk one at a time to identify which specific item caused the exception. In skip-scan mode, each item is run through the full ItemProcessor → ItemWriter pipeline individually. This means ItemProcessor.process() is re-invoked for every item in the chunk — not just the item that caused the exception.
// BillingItemProcessor.java — unsafe mode 3 (UUID computed in process())
@Component
@StepScope
public class BillingItemProcessor
implements ItemProcessor<CustomerBillingItem, CustomerBillingItem> {
@Override
public CustomerBillingItem process(CustomerBillingItem item) throws Exception {
// Developer's reasoning:
// "UUID is computed in the processor, stored in the output object,
// and passed to the writer. The writer uses item.getIdempotencyKey() —
// it doesn't call UUID.randomUUID() itself. The key is stable across
// any write() retries because the processor ran before the retry loop."
//
// The reasoning is correct for FaultTolerantStep RETRY.
// It is incorrect for FaultTolerantStep SKIP-SCAN.
//
// In skip-scan mode, Spring Batch re-invokes ItemProcessor.process()
// for each item individually. A new CustomerBillingItem output is created.
// UUID.randomUUID() generates UUID_B for this item.
// If write() was called in the preceding chunk attempt with UUID_A
// (Stripe committed ch_A), the skip-scan write() call uses UUID_B → ch_B.
String idempotencyKey = UUID.randomUUID().toString(); // UUID_B in skip-scan
item.setIdempotencyKey(idempotencyKey);
return item;
}
}
// BillingJobConfig.java — FaultTolerantStep with both retry and skip
@Bean
public Step billingStep(StepBuilderFactory stepBuilderFactory,
ItemReader<CustomerBillingItem> reader,
BillingItemProcessor processor,
BillingItemWriter writer) {
return stepBuilderFactory.get("billingStep")
.<CustomerBillingItem, CustomerBillingItem>chunk(50)
.reader(reader)
.processor(processor)
.writer(writer)
.faultTolerant()
.retry(StripeConnectException.class)
.retryLimit(3)
.skip(DatabaseConstraintViolationException.class) // DB constraint on billing record
.skipLimit(10)
.build();
}
The skip-scan execution trace
Understanding why this fires requires tracing through the Spring Batch skip-scan algorithm step by step:
- Full-chunk write attempt: The chunk contains 50 items.
ItemProcessor.process()has been called for all 50 items. UUID_A through UUID_50A are stored in the output objects.write(List<50 items>)is called. The writer iterates over items 1 through 38 successfully — Stripe commits pi_1A through pi_38A. On item 39, aDatabaseConstraintViolationExceptionis thrown when attempting to insert the billing record (a duplicate record already exists from a previous partial run). This is a skippable exception — not a retryable one. Spring Batch catches it. - Skip-scan mode initiated: Spring Batch cannot know from the full-chunk write() invocation which item caused the
DatabaseConstraintViolationException. It enters skip-scan mode: reprocesses all 50 items one at a time. - ItemProcessor re-invoked for all 50 items: For each item, Spring Batch calls
processor.process(item)again.UUID.randomUUID()generates UUID_B for each item. Items 1 through 38 now have UUID_B instead of UUID_A. - ItemWriter.write(List<1 item>) called for each re-processed item: For item 1,
write()is called with UUID_B. Stripe looks up UUID_B — it has no record of UUID_B (only UUID_A). Stripe processes a new charge: ch_B for customer 1 alongside the already-committed ch_A. - Final state for customers 1–38: Two Stripe charges each. pi_A committed in the full-chunk attempt. pi_B committed in the skip-scan single-item write. The database has one record per customer (the one successfully inserted during the skip-scan). The application considers the job partially successful (item 39 was skipped). Customers 1–38 have been double-charged. No error was logged for them.
The developer’s reasoning — “UUID is in the processor, not the writer” — is correct for retry mode, where ItemProcessor is not re-invoked. It is incorrect for skip-scan mode, where ItemProcessor is re-invoked for every item in the failed chunk. The distinction between retry and skip-scan in Spring Batch’s fault-tolerance model is non-obvious: both involve re-processing the same chunk in response to an exception, but they differ in whether ItemProcessor is re-invoked.
Why skip-scan re-invokes ItemProcessor
Spring Batch’s skip-scan algorithm needs to identify the individual item that caused the exception when a full-chunk write() call fails. It cannot identify this from the exception alone — the exception was thrown inside write(), which received a list of items. The algorithm must isolate the failing item. It does this by processing each item individually through the full pipeline and observing which single-item write throws the same exception. This requires re-running ItemProcessor because the processor output is what ItemWriter receives — and the processor output from the full-chunk attempt may no longer be available in memory (Spring Batch does not guarantee that the processor output list from the failed attempt is retained for the skip-scan phase).
The fix: content-hash key in the input object, not computed in process()
The safe location for key computation in a FaultTolerantStep with both retry and skip is in the ItemReader or in the input object itself, before either ItemProcessor or ItemWriter is invoked. The ItemReader is not re-invoked in either retry or skip-scan mode — it has already provided the items for the chunk. A key computed during the read phase and embedded in the CustomerBillingItem is stable across all processing and writing attempts, regardless of which fault-tolerance mechanism Spring Batch applies.
// BillingItemReader.java — key computed during read (stable across retry and skip-scan)
@Bean
@StepScope
public JdbcCursorItemReader<CustomerBillingItem> billingItemReader(DataSource dataSource) {
return new JdbcCursorItemReaderBuilder<CustomerBillingItem>()
.name("billingItemReader")
.dataSource(dataSource)
.sql("""
SELECT c.id, c.stripe_customer_id, c.amount_cents,
SHA2(CONCAT(c.id, '|', ?, '|', c.amount_cents), 256) AS idempotency_key
FROM customers c
WHERE c.active = true
""")
.preparedStatementSetter(ps -> ps.setString(1, billingPeriod))
.rowMapper((rs, rowNum) -> {
CustomerBillingItem item = new CustomerBillingItem();
item.setCustomerId(rs.getString("id"));
item.setStripeCustomerId(rs.getString("stripe_customer_id"));
item.setAmountCents(rs.getLong("amount_cents"));
// Key computed by the database during read — before retry or skip-scan.
// Same key in every execution for the same customer + period + amount.
item.setIdempotencyKey(rs.getString("idempotency_key"));
return item;
})
.build();
}
Computing the key in SQL has additional advantages in a Spring Batch context: the key is deterministic across all executions, including restarts and new JobInstances for the same billing period. Even if the Mode 2 failure fires — a new JobInstance re-reads all customers — the key is the same as in the previous execution. Stripe deduplicates all requests for already-charged customers regardless of which execution sent the original request.
For applications where computing the hash in SQL is impractical (non-SQL readers, NoSQL data sources, stream readers), the alternative is to compute the key in the ItemReader’s read() method and embed it in the returned item object:
// BillingItemReader.java — key computed in read() for non-SQL sources
@Component
@StepScope
public class BillingItemReader implements ItemReader<CustomerBillingItem> {
private final CustomerRepository customerRepo;
private final String billingPeriod;
private Iterator<Customer> cursor;
@Override
public CustomerBillingItem read() throws Exception {
if (cursor == null) {
cursor = customerRepo.findAllActive().iterator();
}
if (!cursor.hasNext()) return null;
Customer customer = cursor.next();
CustomerBillingItem item = new CustomerBillingItem(customer);
// Key computed once per item, in the reader — before retry or skip-scan.
item.setIdempotencyKey(contentHashKey(
customer.getId(), billingPeriod, customer.getAmountCents()));
return item;
}
}
Testing Mode 3 with Spring Batch Test and WireMock
Spring Batch’s @SpringBatchTest and JobLauncherTestUtils provide infrastructure for testing the skip-scan path. The test needs to configure WireMock to succeed for items 1–38 and fail with a non-retryable exception for item 39, then verify that items 1–38 each received exactly one Stripe request with a stable key.
// BillingStepSkipScanTest.java — verify skip-scan does not generate UUID_B for items 1-38
@SpringBatchTest
@SpringBootTest
@ExtendWith(WireMockExtension.class)
class BillingStepSkipScanTest {
@Autowired private JobLauncherTestUtils jobLauncherTestUtils;
@Autowired private JobRepositoryTestUtils jobRepositoryTestUtils;
@RegisterExtension
static WireMockExtension wireMock = WireMockExtension.newInstance()
.options(wireMockConfig().port(8089))
.build();
@BeforeEach
void setUp() {
jobRepositoryTestUtils.removeJobExecutions();
}
@Test
void skipScanDoesNotDuplicateChargesForSuccessfulItems() throws Exception {
// Items 1-38: Stripe succeeds
wireMock.stubFor(post(urlPathEqualTo("/v1/payment_intents"))
.withHeader("Idempotency-Key", matching("^(?!.*item-39).*"))
.willReturn(aResponse()
.withStatus(200)
.withBody("""{"id":"pi_A","status":"succeeded","amount":5000}""")));
// Item 39: DB constraint violation on write — triggers skip-scan
// (simulated by configuring the test DB with a pre-existing record for item 39)
// Stripe stub for item 39 in skip-scan: should only be called once
wireMock.stubFor(post(urlPathEqualTo("/v1/payment_intents"))
.withHeader("Idempotency-Key", containing("item-39"))
.willReturn(aResponse()
.withStatus(200)
.withBody("""{"id":"pi_39","status":"succeeded","amount":5000}""")));
JobExecution execution = jobLauncherTestUtils.launchStep("billingStep");
// Job should complete with COMPLETED status (item 39 was skipped)
assertEquals(BatchStatus.COMPLETED, execution.getStatus());
assertEquals(1, execution.getStepExecutions().iterator().next().getSkipCount());
// Verify items 1-38 each received exactly ONE Stripe request
List<LoggedRequest> requests = wireMock.findAll(
postRequestedFor(urlPathEqualTo("/v1/payment_intents")));
// Group by Idempotency-Key — each key should appear exactly once
Map<String, Long> keyFrequency = requests.stream()
.collect(Collectors.groupingBy(
r -> r.getHeader("Idempotency-Key"),
Collectors.counting()));
// With content-hash keys: every key appears exactly once (Stripe deduplicates)
// With UUID.randomUUID() in process(): keys from the chunk attempt and skip-scan
// differ — two requests per item with different keys — ch_B alongside ch_A
keyFrequency.forEach((key, count) -> assertEquals(1L, count,
"Key " + key + " appeared " + count + " times — duplicate Stripe request detected"));
}
}
This test structure reveals the Mode 3 failure: with UUID.randomUUID() in ItemProcessor.process(), the assertion fails because each key appears twice — once from the full-chunk attempt and once from the skip-scan single-item write, with different keys for the same item. With content-hash keys, the keys from both phases are identical, Stripe deduplicates on the second call, and each key appears exactly once.
Comparison: the three Spring Batch Stripe failure modes
| Failure mode | Re-execution unit | What re-invokes | UUID position (unsafe) | Developer misconception |
|---|---|---|---|---|
| Mode 1: FaultTolerantStep retry | Entire write(List<N items>) method |
RetryTemplate wrapping write() |
Inside write() body, per-item loop |
“Spring Batch retries the chunk, not write() — same invocation re-attempted” |
| Mode 2: @Scheduled + new JobInstance | Entire step (all items from read start) | @Scheduled triggers jobLauncher.run() with new run.id |
Inside write() or process() |
“Adding run.id is the restart pattern — new run picks up where old one left off” |
| Mode 3: Skip-scan re-processing | Each item individually through full pipeline | Skip-scan algorithm re-invokes process() + write() |
Inside process() |
“UUID is in the processor, not the writer — stable across write() retries” |
| Failure mode | Stripe impact | DB impact | Detection signal |
|---|---|---|---|
| Mode 1: Retry | ch_B for every item in chunk (not just the failed item) | Duplicate DB records if write() also inserts | Two Stripe charges per customer, retry log entry in Spring Batch metadata |
| Mode 2: New JobInstance | ch_B for all items processed in failed first run | Duplicate DB records or constraint violations | Two Stripe charges per customer, two JobInstances in BATCH_JOB_INSTANCE table for same period |
| Mode 3: Skip-scan | ch_B for all items in failed chunk (except the skipped item) | One DB record per item (skip-scan single-item write succeeded) | Two Stripe charges per customer, one DB record, skip count = 1 in step execution |
Querying Spring Batch metadata tables to detect which mode fired
Spring Batch persists execution metadata in tables that reveal which failure mode occurred. Querying these tables is the fastest way to determine whether duplicate charges exist and how many customers are affected.
-- Detect Mode 2: multiple JobInstances for the same billing period
-- Each row represents a distinct JobInstance — multiple rows = new JobInstance per run
SELECT ji.JOB_INSTANCE_ID, ji.JOB_NAME, jep.KEY_NAME, jep.STRING_VAL,
je.START_TIME, je.END_TIME, je.STATUS, je.EXIT_CODE
FROM BATCH_JOB_INSTANCE ji
JOIN BATCH_JOB_EXECUTION je ON ji.JOB_INSTANCE_ID = je.JOB_INSTANCE_ID
JOIN BATCH_JOB_EXECUTION_PARAMS jep ON je.JOB_EXECUTION_ID = jep.JOB_EXECUTION_ID
WHERE ji.JOB_NAME = 'billingJob'
AND jep.KEY_NAME = 'billingPeriod'
AND jep.STRING_VAL = '2026-10'
ORDER BY je.START_TIME;
-- If multiple rows: each started from scratch — all items processed multiple times
-- Detect Mode 1: step executions with retry count > 0 on a billingStep
SELECT se.STEP_NAME, se.START_TIME, se.STATUS,
se.COMMIT_COUNT, se.ROLLBACK_COUNT, se.READ_COUNT, se.WRITE_COUNT
FROM BATCH_STEP_EXECUTION se
JOIN BATCH_JOB_EXECUTION je ON se.JOB_EXECUTION_ID = je.JOB_EXECUTION_ID
WHERE se.STEP_NAME = 'billingStep'
AND se.ROLLBACK_COUNT > 0
ORDER BY se.START_TIME;
-- ROLLBACK_COUNT > 0 indicates at least one chunk required retry or rollback
-- Detect Mode 3: step executions with skip count > 0
SELECT se.STEP_NAME, se.START_TIME, se.COMMIT_COUNT, se.SKIP_COUNT,
se.PROCESS_SKIP_COUNT, se.WRITE_SKIP_COUNT
FROM BATCH_STEP_EXECUTION se
WHERE se.STEP_NAME = 'billingStep'
AND (se.SKIP_COUNT > 0 OR se.WRITE_SKIP_COUNT > 0);
-- WRITE_SKIP_COUNT > 0 indicates skip-scan was triggered for at least one item
For Mode 3, correlating the WRITE_SKIP_COUNT with Stripe’s payment intent list for the affected billing period reveals the actual blast radius. Stripe’s PaymentIntent.list() with filters for customer and creation date returns all charges. Two charges with consecutive creation times and different IDs for the same customer indicate a Mode 3 duplicate. The skip_count in the step execution tells you the chunk where skip-scan fired — multiplied by the chunk size, this gives you the maximum number of customers potentially double-charged.
Kotlin Spring Batch: the same three modes with coroutine readers
Spring Batch’s Kotlin DSL, introduced in Spring Batch 5, uses the same underlying fault-tolerance infrastructure as the Java API. The step DSL wraps the same FaultTolerantStep builder. The retry and skip-scan mechanics are identical. Kotlin’s val scoping does not provide any protection: a val inside a function body is a new binding on every invocation of that function, including retry invocations by RetryTemplate.
// BillingJobConfig.kt — Kotlin Spring Batch 5 DSL (same fault-tolerance mechanics)
@Bean
fun billingStep(jobRepository: JobRepository, transactionManager: PlatformTransactionManager,
reader: ItemReader<CustomerBillingItem>,
processor: ItemProcessor<CustomerBillingItem, CustomerBillingItem>,
writer: ItemWriter<CustomerBillingItem>) = step("billingStep") {
chunk<CustomerBillingItem, CustomerBillingItem>(50, transactionManager) {
reader(reader)
processor(processor)
writer(writer)
faultTolerant {
retry<StripeConnectException> { limit(3) }
skip<DatabaseConstraintViolationException> { limit(10) }
}
}
}
// BillingItemProcessor.kt — unsafe (UUID in process() — Kotlin val doesn't help)
@Component
@StepScope
class BillingItemProcessor : ItemProcessor<CustomerBillingItem, CustomerBillingItem> {
override fun process(item: CustomerBillingItem): CustomerBillingItem {
// val is immutable within this invocation — not across invocations.
// Skip-scan re-invokes process() — new invocation = new val binding = UUID_B.
val idempotencyKey = UUID.randomUUID().toString() // UUID_B in skip-scan
return item.copy(idempotencyKey = idempotencyKey)
}
}
// BillingItemProcessor.kt — fixed (content-hash key stable across all re-invocations)
@Component
@StepScope
class BillingItemProcessor(
@Value("#{jobParameters['billingPeriod']}") private val billingPeriod: String
) : ItemProcessor<CustomerBillingItem, CustomerBillingItem> {
override fun process(item: CustomerBillingItem): CustomerBillingItem {
// Content-hash key: same inputs = same key, in every invocation.
// Stable across FaultTolerantStep retry, skip-scan, and new JobInstance re-runs.
val idempotencyKey = contentHashKey(item.customerId, billingPeriod, item.amountCents)
return item.copy(idempotencyKey = idempotencyKey)
}
private fun contentHashKey(customerId: String, period: String, amountCents: Long): String {
val input = "$customerId|$period|$amountCents"
return MessageDigest.getInstance("SHA-256")
.digest(input.toByteArray(Charsets.UTF_8))
.joinToString("") { "%02x".format(it) }
}
}
Injecting billingPeriod via @Value("#{jobParameters['billingPeriod']}") into the @StepScope-scoped processor bean ensures the content-hash computation uses the same billing period value that the JobInstance was launched with. This is correct for fix option A (single JobInstance per period) and catches Mode 2 as well as Mode 3: even if a new JobInstance is launched with a different run.id but the same billingPeriod, the key computation produces the same hash.
The complete fix: key placement rules for Spring Batch Stripe jobs
The three failure modes in this post share a single root cause: idempotency keys are computed inside a code path that Spring Batch’s fault-tolerance infrastructure re-invokes. The fix in all three cases is to move key computation to a location that is not re-invoked by any retry, skip-scan, or restart mechanism.
Key placement rules for Spring Batch Stripe billing jobs:
- Never compute the key inside
ItemWriter.write().write()is re-invoked byRetryTemplateon retry (Mode 1) and by skip-scan (Mode 3). Any random or time-based computation insidewrite()regenerates on every invocation. - Do not compute the key inside
ItemProcessor.process()if the step uses skip.process()is re-invoked by skip-scan (Mode 3) even if it is not re-invoked by retry. If the step uses only retry (no skip),process()is safe for key computation — but a step that adds skip later will silently break this invariant. - The safest location is the
ItemReaderor the input data source.read()is not re-invoked by retry, skip-scan, or restart (Spring Batch advances the reader cursor past committed chunks on restart, but does not re-read already-read items within the current chunk). Keys computed here are stable across all fault-tolerance paths. - Use content-hash keys, not UUID. Content-hash keys dedup across not just retry and skip-scan but also across new
JobInstances for the same logical billing period (Mode 2). A key derived fromcustomerId + billingPeriod + amountCentsis the same in every execution that processes the same charge, regardless of which Spring Batch construct triggered that execution. - Fix the
@Scheduledlaunch pattern separately. Content-hash keys at the Stripe layer prevent double-charging, but they do not prevent your database from receiving duplicate inserts on a newJobInstancere-run. Fix the scheduler to useJobOperator.restart()for failed executions, and audit whether every billing job is truly restartable (step configured withallowStartIfComplete(false)or idempotent writes withON CONFLICT DO NOTHING).
Related posts on retry-triggered Stripe duplicate charges
The Spring Batch failure modes in this post are three instances of a broader pattern: any retry mechanism that re-invokes code containing UUID.randomUUID() generates a new idempotency key and a new Stripe charge. The same pattern appears in:
- Spring WebClient reactive retry —
Mono.retryWhen()re-subscribes to the coldMono.fromCallable()— UUID inside the callable generates UUID_B on re-subscription - Spring @Async + @Transactional —
CompletableFutureretry in a@Scheduledexecutor — UUID inside the async task generates UUID_B on the retry thread - Kotlin Coroutines @Transactional retry — Arrow’s
retry{}and Spring’s@Retryableon suspend functions - Kotlin Flow + Exposed — Exposed’s automatic deadlock retry,
Flow.retry{}re-collection, andflatMapLatest{}cancellation
In every case, the fix is the same: compute the idempotency key before the retry scope, using deterministic inputs that produce the same key for the same logical charge regardless of how many times the surrounding code re-executes.
Keybrake: enforce idempotency at the proxy layer regardless of application code
All three failure modes in this post are variants of the same root cause: application code generates a new idempotency key when Spring Batch’s fault-tolerance infrastructure re-invokes a billing code path. The retry mechanism differs — RetryTemplate, a new JobInstance, skip-scan re-processing — but the outcome is the same: UUID_B sent to Stripe, ch_B created alongside ch_A.
Content-hash keys fix the application-layer root cause for the Stripe calls themselves. But in a system where multiple batch jobs, multiple developers, and multiple versions of the batch configuration exist, enforcing the “compute the key before the fault-tolerance scope” rule across every billing job is an auditing problem. New jobs ship without review; content-hash keys get accidentally replaced with UUID.randomUUID() in a processor refactor; a new skip configuration is added to a step without realizing the processor now needs a deterministic key.
Keybrake is a scoped API-key proxy that sits between your application and Stripe. When you route your Stripe calls through Keybrake instead of directly to api.stripe.com, Keybrake can enforce idempotency at the proxy layer: it normalizes idempotency keys to deterministic hashes based on the request body, detects when a key changes across calls for the same logical charge, and either deduplicates or alerts — depending on your policy. This catches the class of bug described in this post regardless of which Spring Batch fault-tolerance mechanism triggers the re-invocation, without requiring code changes in every batch job that calls Stripe.
Put the brakes on your agent’s Stripe calls
Join the waitlist for Keybrake — scoped API keys, spend caps, and audit log for the SaaS APIs your agents call.