What it says
What this says
Current to 11 Oct 26Final internal review dated 29 Jul 2026 of a nightly refresh failure after release four-twelve. Brains stayed available but answered on stale data, a freshness failure. Two main actions were open at review; the cause was a scheduler change plus no completeness alerting.
What it is: BB-Demo's own incident report, written by Ewan Pryce and Kieran Walsh. It says "That is a freshness failure, not an availability failure." Every customer on every tier was affected.
Cause: the incremental-load scheduler change in release four-twelve let the refresh finish "without loading everything it should have". Contributing factors were no completeness check and a junior engineer on call with one senior platform engineer.
Commits: no customer-facing service credit is stated. It says "Only Enterprise is credited for data freshness" and that consequences are handled account by account by customer success.
Not settled: whether customers acted on stale answers, whether other scheduler changes carry similar risk, and which accounts are owed credits. The document gives no dates for the actions and both main actions were open at review.
As found in this document
Current to 9 Oct 26- Release coverNo release without a senior platform engineer on call. A release waits until one is named as cover.No release without a senior platform engineer on call. A release waits until one is named as cover. · Source: object_read:doc_co_016
- Refresh-completeness alertingalert when the nightly refresh finishes but loaded less than expected for any connector, not only when it fails.alert when the nightly refresh finishes but loaded less than expected for any connector, not only when it fails. · Source: object_read:doc_co_016
- Refresh targetThe refresh had not finished by its usual 07:00 target.The refresh had not finished by its usual 07:00 target. · Source: object_read:doc_co_016
- Freshness creditOnly Enterprise is credited for data freshnessOnly Enterprise is credited for data freshness · Source: object_read:doc_co_016
The document
Body
BB-Demo - Incident Report
Status: Final Review date: Wednesday 29 July 2026 Authors: Ewan Pryce, Head of Platform; Kieran Walsh, Platform Engineer Subject: July nightly refresh failure following release four-twelve
Summary
In short: release four-twelve went out at 18:30 and included a change to the incremental-load scheduler. After that release the nightly refresh did not complete as it should have. Customers' brains stayed available and answered questions, but on data that had not been refreshed. That is a freshness failure, not an availability failure.
This review sets out what we know, what we do not, and what we are changing. It is written without blame. The release process allowed this to happen, and the process is what we are fixing.
Impact
Read the whole document (5,525 characters)
- Who was affected: customers whose brains depend on the nightly refresh, which is every customer on every tier.
- What they saw: chat, search and pages responded normally. The data behind the answers was older than it should have been, so teams starting their day could be looking at yesterday's picture without knowing it.
- Why it was easy to miss: the platform was up the whole time. Nothing on our side reported an outage, because availability monitoring checks that the brain answers, not that the data behind it is current.
- Service levels: availability and freshness are separate commitments. A brain answering on stale data was available. Only Enterprise is credited for data freshness, so the commercial consequence differs by tier and is being handled account by account by customer success.
Timeline
| Stage | What happened |
|---|---|
| Release evening, 18:30 | Release four-twelve deployed, including the incremental-load scheduler change. |
| Overnight | The scheduled nightly refresh ran under the new scheduler and did not complete across all connectors. |
| Following morning | The refresh had not finished by its usual 07:00 target. Customers' first questions of the day were answered from stale data. |
| Following working hours | Platform team confirmed the refresh was incomplete, raised a ticket for each affected area and began diagnosis. |
| Recovery | The scheduler change was dealt with and the refresh re-run, and completeness was checked by hand connector by connector before we called it done. |
The on-call engineer for the release was Kieran, in his seventh week at BB-Demo. He raised the issue promptly and kept a clear log, which made this review much easier to write.
Cause
Root cause, not symptom: the incremental-load scheduler change shipped in release four-twelve altered how the scheduler decided what each connector should load. Under that logic the refresh could finish its run without loading everything it should have, and nothing in our monitoring noticed the difference between "ran" and "completed".
There are two contributing factors, and we would rather name them than hide behind the first.
- No completeness check. We alert when the refresh job fails. We do not alert when the job succeeds but loads less than expected. A refresh that finishes quietly and incompletely looks identical to a good one.
- Release cover. The release went out in the evening with a junior engineer on call and no senior platform engineer available to take the first call. Since the departure of a senior colleague in May, the platform team has had one senior engineer. That is a structural weakness, and it is on us.
What went well
- The problem was identified the same working day it was visible.
- Nobody said a fix was done until it had been checked against the data.
- Customer success had accurate, plain information to take to customers.
What did not go well
- We found out from the morning rather than from an alert.
- The release was scheduled at a time when the on-call arrangement could not absorb a problem of this kind.
- The release notes described the scheduler change accurately but we had not agreed with ourselves what "good" looked like on the first morning, so there was nothing to compare against.
Actions
| Action | Owner | Status |
|---|---|---|
| Refresh-completeness alerting: alert when the nightly refresh finishes but loaded less than expected for any connector, not only when it fails. | Ewan Pryce | Open |
| No release without a senior platform engineer on call. A release waits until one is named as cover. | Ewan Pryce | Open |
| Agree and write down, before each release that touches the refresh, what a healthy first morning looks like. | Kieran Walsh, with Ewan Pryce | Proposed |
Both main actions were open at the time of this review. Neither is done until it has been shown working, and we will report on them in the same way we have reported on this incident.
Customer communication
Customer success are speaking to affected customers directly. The line they have is the one in this report: the brain was available, the data was stale, the cause was ours, and we are putting alerting and release cover in place. They have been asked not to promise more than that.
Not yet known
- Whether any customer made a decision in the affected window on the basis of stale answers. We will only find that out by asking, and customer success are asking.
- Whether other scheduler changes in the pipeline carry a similar risk. We are checking before anything else touching the refresh is released.
Sign-off
Ewan Pryce, Head of Platform Kieran Walsh, Platform Engineer
Reviewed Wednesday 29 July 2026
Ewan
Unusual terms
Current to 9 Oct 26- Freshness failure without an availability breach
- Only Enterprise credited for freshness