> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tokenfactory.nebius.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Run concurrent workloads

> Bound parallel operations, correlate state, classify failures, and cancel only work owned by the batch.

Use a semaphore or fixed worker pool. Give each work item a stable application ID and record:

| Field | Purpose |
| - | - |
| Item ID and attempt | Application correlation and duplicate detection |
| Input image UUID | Reproducible starting state |
| Operation ID | Status, result, and cancellation target |
| Persistence mode | Whether a result image is expected |
| Terminal status | Platform outcome: `SUCCESS`, `FAILED`, or `CANCELLED` |
| Process state | Exit code, signal, or timed-out flag |
| Result image UUID | Continuation point when present |

Tags can move while a batch runs; resolve them before scheduling. Plan for no more than 50 simultaneous operations, then keep the application bound below the effective limit returned for the current token. Leave room for recovery or interactive work.

## Classify before retrying

| Observation | Safe action |
| - | - |
| Request validation or authorization rejection | Correct the request; no accepted operation ID means no status target |
| Rate or service rejection | Back off according to the response and inspect token limits; do not label an undocumented 503 as capacity |
| Transport failure before an operation ID | Acceptance is uncertain; use application correlation and avoid blind resubmission |
| Client wait or observation timeout with an ID | Query that operation before resubmitting |
| `FAILED` platform status | Inspect platform error; retry from known input only when safe |
| `SUCCESS` plus timed-out flag, signal, or nonzero exit | Classify as workload timeout, signal, or failure |
| Result collection failure | Retry retrieval by recorded operation/image ID |

Create requests have no documented idempotency key. An application item ID helps reconcile duplicates but does not make execution idempotent.

## Stop safely

On an early-stop condition, stop scheduling, cancel unfinished operation IDs owned by this batch with `DELETE /v1/operations/{operation_id}`, and observe each terminal state until a deadline. Never use a project-wide cancel-all action from shared automation. Resume cancelled work from its recorded input image UUID. Follow the [result-image rules](/sandboxes/concepts/images-checkpoints-and-branches#interpret-the-result-image) when recovering files from cancelled work. Contact the Sandboxes team if cancellation retention or usage affects capacity or cost controls.

See [Run a bounded evaluation batch](/sandboxes/cookbook/run-bounded-evaluation-batch), [Troubleshooting](/sandboxes/operate/troubleshooting), and [Limits, retention, and usage](/sandboxes/operate/limits-retention-and-usage).
