Snowflake Ends Broad Gen1 Advice; SPCS Remains Constrained
Warehouse capacity improved. SPCS remains constrained.
Snowflake reports improved warehouse capacity in GCP US Central1. The incident began 9 August at 08:41 UTC. The guidance changed 8 October at 21:37:51 UTC, and status remains Identified. No customer count or independent performance measurement is published.
The operational mistake would be to turn that mixed state into one estate-wide answer. A warehouse that starts normally and a container workload waiting for new capacity are different tests, with different failure evidence and different rollback choices.
Split the recovery inventory
Create two queues before changing anything. Put standard virtual warehouses in one and Snowpark Container Services workloads in the other. For each entry, record the account, region, owner, workload, present configuration, failed operation and the evidence behind any temporary measure.
For a warehouse, compare start and resume success, time to useful capacity, representative query completion and pipeline deadlines against a recent healthy baseline. Run the test during the demand conditions that previously mattered. A successful quiet-period start is useful evidence, but it does not answer the peak-demand question.
For SPCS, preserve the request identifier, compute pool, instance family, service or job, timestamp, error and current state. Check whether a failed request created capacity or triggered work before deciding that it is safe to retry.
Review Gen1 as a controlled exception
Snowflake ended broad Gen1 advice. Customers who switched can assess Gen2 with Support; Gen1 remains an option for verified reliability problems.
Snowflake’s Gen2 documentation says standard warehouses can change generation and advises workload-specific cost and performance testing. Treat that capability as a change mechanism, not a decision.
BlackTree recommends this review sequence:
- Identify only the warehouses changed because of this incident.
- Ask Support to confirm the proposed assessment window and any workload-specific constraint.
- Select one low-risk representative warehouse.
- Define success, failure and rollback conditions before the assessment.
- Compare start and resume behaviour, query results, pipeline completion, performance and cost.
- Expand only when the same evidence is repeatable. Keep documented exceptions where it is not.
This process does not authorise a universal generation change. The account configuration, workload and agreed support path still control the decision.
Reconcile SPCS work before retrying
SPCS work needing new capacity can still fail or wait. Warehouse Notebooks may defer SPCS migration. There is no general workaround.
Snowflake defines a compute pool as the virtual-machine nodes that run SPCS services and jobs. That boundary makes the retry record concrete: identify the pool and operation, then check the result and any external side effect.
Use an idempotency check for job submission, service creation and automation around provisioning. If the first result is ambiguous, preserve the evidence and ask Support to reconcile it. Repeating an uncertain operation can create duplicate work without adding useful recovery evidence.
No universal region failover follows from this incident. A different region introduces data, network, identity, latency and support decisions that need their own approved recovery design.
Close on evidence, not the calendar
Snowflake expects added capacity by late October, but does not promise recovery then.
Define closure before the next test. Useful checks include repeatable warehouse starts and resumes under representative demand, completed pipelines, successful SPCS provisioning where required, reconciled failed jobs and closed support exceptions. Record the workload and observation window behind each result.
BlackTree’s DigitalOcean BLR1 analysis applies the same method to a distinct provider incident: reconcile dependent work rather than assuming that a service-state change completed every request.
No CVE applies. Keep security, data-loss and universal-outage claims out of communications unless separate evidence supports them.
Sources: Snowflake incident, Gen2 warehouse documentation, and SPCS compute-pool documentation.


