apache/seatunnel · warning
WAL send failed for record id=
Error message
WAL send failed for record id={}, batchId={}; row stays SENDING for resurrect/retry What it means
EdgeAgentRuntimeScheduler.sendClaimedRecords catches IOException while sending a WAL record to the remote endpoint. Non-decrypt failures are logged as a warning: the record deliberately stays in SENDING state so a later resurrection/retry pass can resend it. Decrypt failures are rethrown because they indicate a non-transient problem.
Solutions
- Check the logged record id/batchId and exception to confirm it is a transient network issue.
- Verify the receiver endpoint is reachable and healthy (ports, auth, load).
- Wait for the retry/resurrect mechanism to resend the SENDING record; ensure backoff settings are sane.
- If failures persist, restart the scheduler after confirming connectivity; check that decrypt keys/config are correct if the error ever surfaces as decrypt-failed.
Example fix
// before: no retry tuning, short timeout kills long sends
transport {
timeout-ms = 1000
}
// after: tolerate slower links
transport {
timeout-ms = 30000
max-backoff-ms = 60000
} Defensive patterns
Strategy: retry
Validate before calling
// Pre-send health probe
if (!transport.isSessionHealthy()) {
transport.ensureAuthenticatedSession();
} Try / catch
try {
send(record);
} catch (IOException ex) {
if (!isDecryptFailed(ex)) {
// leave record in SENDING; rely on resurrect/retry with capped backoff
backoffSleep(attempt++);
} else {
throw ex; // non-transient: surface immediately
}
} Prevention
- Keep SENDING-state resurrect/retry enabled and monitor retry depth.
- Alert on receiver health so sockets are not dropped mid-batch.
- Set realistic transport timeouts for batch sizes on your network.
- Separate decrypt/config errors (fail fast) from transport errors (retry).
When it happens
Trigger: sendClaimedRecords -> sending a claimed WAL record over the transport throws IOException (connection drop, socket reset, broken pipe) that is not classified as a decrypt failure.
Common situations: Receiver restarted mid-batch, network partition between agent and receiver, receiver overloaded closing sockets, or TLS issues corrupting the stream.
Related errors
- [ ] request http failed
- Edge socket receiver loop exception, retrying
- Edge transport IO failure, will reconnect. batchId=
- Failed to execute HTTP request to Firebase endpoint
- Failed to publish NATS JetStream message for subtask
AI-assisted analysis of apache/seatunnel@cf67b549a7 (2026-09-10).
Data as JSON: /api/errors/24614e0915f4f7a6.
Report an issue: GitHub.
Appendix: source
Thrown at seatunnel-edge-agent/seatunnel-edge-agent-starter/src/main/java/org/apache/seatunnel/edge/agent/starter/runtime/EdgeAgentRuntimeScheduler.java:190
int sent = sendClaimedRecords();
return appended > 0 || sent > 0 || !events.isEmpty();
}
private int sendClaimedRecords() throws Exception {
List<WalRecord> records = walStore.claimPending(maxPollRecords, maxAttempts);
int sent = 0;
for (WalRecord record : records) {
long batchId = record.getBatchId() > 0 ? record.getBatchId() : record.getId();
try {
transport.send(batchId, payloadSerializer.serialize(record.getPayload()));
walStore.ack(record.getId());
sent++;
sendFailureAttempt = 0;
} catch (IOException ex) {
if (isDecryptFailed(ex)) {
throw ex;
}
LOG.warn(
"WAL send failed for record id={}, batchId={}; row stays SENDING for"
+ " resurrect/retry",
record.getId(),
batchId,
ex);
Thread.sleep(
EdgeTransportConfig.computeBackoffMillis(
sendFailureAttempt++, backoffMs, backoffMaxMs));
break;
} catch (InterruptedException ex) {
Thread.currentThread().interrupt();
throw ex;
}
}
return sent;
}
private static boolean isDecryptFailed(IOException ex) {View on GitHub (pinned to cf67b549a7)