LLM decode — generated text, two clouds

Qwen2.5-1.5B-Instruct, fp16, 56 prompt tokens → 256 new tokens, greedy decode, same image digest on both clouds.

Both clouds produced byte-identical text — expected from the same model and deterministic decode; the excerpt below is capped at 400 characters by the job itself (jobs/llm/run.py), not by log capture. What differs between clouds is speed, not output.

Azure · Tesla T4 ↑ winner
36.1 tok/s (samples 36.1, 36.1, 36.1)
For engineers working with data or performance metrics, it's important to understand the difference between using a single statistic like the mean (average) versus looking at individual items in a dataset. ### Mean vs. Individual Item Listing #### Mean: - **Definition**: The mean is calculated by summing all values and dividing by the number of values. - **Example**: If you have five numbers: 10 [cut at 400 chars]
cloud=azure · job=llm · execution=llm-gpu-0dn96x9 · 2026-08-29 · commit bf8a95f · model_load 27.9s
GCP · NVIDIA L4
29.7 tok/s (samples 29.4, 29.7, 29.7)
For engineers working with data or performance metrics, it's important to understand the difference between using a single statistic like the mean (average) versus looking at individual items in a dataset. ### Mean vs. Individual Item Listing #### Mean: - **Definition**: The mean is calculated by summing all values and dividing by the number of values. - **Example**: If you have five numbers: 10 [cut at 400 chars]
cloud=gcp · job=llm · execution=llm-gpu-jnktw (Cloud Run europe-west4) · 2026-09-02 · commit bf8a95f · model_load 18.7s