d2l-ai/d2l-zh · error · AssertionError
train_acc <= 1 and train_acc > 0.7
Error message
train_acc <= 1 and train_acc > 0.7
What it means
An AssertionError in d2l.paddle.train_ch3 requiring final training accuracy in (0.7, 1]. It is a post-training sanity check for the PaddlePaddle softmax-regression benchmark on Fashion-MNIST; accuracy <= 0.7 signals undertraining or divergence, while > 1 signals a broken metric computation.
Source
Thrown at d2l/paddle.py:351
for x, y, fmt in zip(self.X, self.Y, self.fmts):
self.axes[0].plot(x, y, fmt)
self.config_axes()
display.display(self.fig)
display.clear_output(wait=True)
def train_ch3(net, train_iter, test_iter, loss, num_epochs, updater):
"""训练模型(定义见第3章)
Defined in :numref:`sec_softmax_scratch`"""
animator = Animator(xlabel='epoch', xlim=[1, num_epochs], ylim=[0.3, 0.9],
legend=['train loss', 'train acc', 'test acc'])
for epoch in range(num_epochs):
train_metrics = train_epoch_ch3(net, train_iter, loss, updater)
test_acc = evaluate_accuracy(net, test_iter)
animator.add(epoch + 1, train_metrics + (test_acc,))
train_loss, train_acc = train_metrics
assert train_loss < 0.5, train_loss
assert train_acc <= 1 and train_acc > 0.7, train_acc
assert test_acc <= 1 and test_acc > 0.7, test_acc
def predict_ch3(net, test_iter, n=6):
"""预测标签(定义见第3章)
Defined in :numref:`sec_softmax_scratch`"""
for X, y in test_iter:
break
trues = d2l.get_fashion_mnist_labels(y)
preds = d2l.get_fashion_mnist_labels(d2l.argmax(net(X), axis=1))
titles = [true +'\n' + pred for true, pred in zip(trues, preds)]
d2l.show_images(
d2l.reshape(X[0:n], (n, 28, 28)), 1, n, titles=titles[0:n])
def evaluate_loss(net, data_iter, loss):
"""评估给定数据集上模型的损失。
Defined in :numref:`sec_model_selection`"""View on GitHub (pinned to e6b18ccea7)
Solutions
- Train the full 10 epochs with lr=0.1, batch_size=256.
- Audit train_epoch_ch3's metric.Accumulator(3): add correct-count and count per batch and divide metric[1] by metric[2] exactly once.
- Verify optimizer.step() / linear.clear_gradients() (or paddle.optimizer SGD minimization) run every batch so weights actually update.
- If acc hovers near 0.1, debug data loading (labels aligned with images) before touching hyperparameters.
Example fix
# before train_ch3(net, train_iter, test_iter, loss, 1, updater) # acc <= 0.7 # after train_ch3(net, train_iter, test_iter, loss, 10, updater)
Defensive patterns
Strategy: validation
Validate before calling
train_loss, train_acc = train_epoch_ch3(net, train_iter, loss, updater)
if not (0.7 < train_acc <= 1.0):
raise ValueError(f'train_acc={train_acc:.3f} outside (0.7, 1]; '
f'check optimizer.step()/clear_gradients and metric Accumulator') Try / catch
try:
train_ch3(net, train_iter, test_iter, loss, num_epochs, updater)
except AssertionError as e:
raise RuntimeError(f'Paddle train_ch3 accuracy check failed ({e}); '
f'likely undertrained or metrics mis-accumulated') from e Prevention
- Ensure optimizer.step() and gradient clearing run every batch.
- Keep the Accumulator pattern: add(correct_count, count) and divide once at the end.
- Train the full 10 epochs; near-chance accuracy (~0.1) means updates are not applied.
When it happens
Trigger: Undertrained model (1-2 epochs instead of 10); wrong gradient scaling in a custom updater producing near-chance accuracy; train_epoch_ch3 accuracy accumulator dividing by the wrong count (giving acc > 1); shuffled label pipeline making learning impossible.
Common situations: Cutting num_epochs short when iterating quickly; adapting train_epoch_ch3 to new metrics and breaking the Accumulator bookkeeping; Paddle dtype/device mismatches silently zeroing gradients.
Related errors
- train_acc <= 1 and train_acc > 0.7
- train_acc <= 1 and train_acc > 0.7
- train_acc <= 1 and train_acc > 0.7
- train_loss < 0.5
- test_acc <= 1 and test_acc > 0.7
AI-assisted analysis of d2l-ai/d2l-zh@e6b18ccea7 (2026-08-14).
Data as JSON: /api/errors/db4d03ded1f23746.
Report an issue: GitHub.