B-Brain - Incident Report

Documentfinal

incident_report · final · 2026-07-29

Final internal review dated 29 Jul 2026 of a nightly refresh failure after release four-twelve. Brains stayed available but answered on stale data, a freshness failure. Two main actions were open at review; the cause was a scheduler change plus no completeness alerting.

What it says

What this says

Current to 11 Oct 26

Final internal review dated 29 Jul 2026 of a nightly refresh failure after release four-twelve. Brains stayed available but answered on stale data, a freshness failure. Two main actions were open at review; the cause was a scheduler change plus no completeness alerting.

What it is: B-Brain's own incident report, written by Ewan Pryce and Kieran Walsh. It says "That is a freshness failure, not an availability failure." Every customer on every tier was affected.

Cause: the incremental-load scheduler change in release four-twelve let the refresh finish "without loading everything it should have". Contributing factors were no completeness check and a junior engineer on call with one senior platform engineer.

Commits: no customer-facing service credit is stated. It says "Only Enterprise is credited for data freshness" and that consequences are handled account by account by customer success.

Not settled: whether customers acted on stale answers, whether other scheduler changes carry similar risk, and which accounts are owed credits. The document gives no dates for the actions and both main actions were open at review.

AI · claude-sonnet-5-5 · 11 Oct 2026

As found in this document

Current to 9 Oct 26
  • Release coverNo release without a senior platform engineer on call. A release waits until one is named as cover.No release without a senior platform engineer on call. A release waits until one is named as cover. · Source: object_read:doc_co_016
  • Refresh-completeness alertingalert when the nightly refresh finishes but loaded less than expected for any connector, not only when it fails.alert when the nightly refresh finishes but loaded less than expected for any connector, not only when it fails. · Source: object_read:doc_co_016
  • Refresh targetThe refresh had not finished by its usual 07:00 target.The refresh had not finished by its usual 07:00 target. · Source: object_read:doc_co_016
  • Freshness creditOnly Enterprise is credited for data freshnessOnly Enterprise is credited for data freshness · Source: object_read:doc_co_016

The document

Body

B-Brain - Incident Report

Status: Final Review date: Wednesday 29 July 2026 Authors: Ewan Pryce, Head of Platform; Kieran Walsh, Platform Engineer Subject: July nightly refresh failure following release four-twelve

Summary

In short: release four-twelve went out at 18:30 and included a change to the incremental-load scheduler. After that release the nightly refresh did not complete as it should have. Customers' brains stayed available and answered questions, but on data that had not been refreshed. That is a freshness failure, not an availability failure.

This review sets out what we know, what we do not, and what we are changing. It is written without blame. The release process allowed this to happen, and the process is what we are fixing.

Impact

Read the whole document (5,525 characters)
  • Who was affected: customers whose brains depend on the nightly refresh, which is every customer on every tier.
  • What they saw: chat, search and pages responded normally. The data behind the answers was older than it should have been, so teams starting their day could be looking at yesterday's picture without knowing it.
  • Why it was easy to miss: the platform was up the whole time. Nothing on our side reported an outage, because availability monitoring checks that the brain answers, not that the data behind it is current.
  • Service levels: availability and freshness are separate commitments. A brain answering on stale data was available. Only Enterprise is credited for data freshness, so the commercial consequence differs by tier and is being handled account by account by customer success.

Timeline

StageWhat happened
Release evening, 18:30Release four-twelve deployed, including the incremental-load scheduler change.
OvernightThe scheduled nightly refresh ran under the new scheduler and did not complete across all connectors.
Following morningThe refresh had not finished by its usual 07:00 target. Customers' first questions of the day were answered from stale data.
Following working hoursPlatform team confirmed the refresh was incomplete, raised a ticket for each affected area and began diagnosis.
RecoveryThe scheduler change was dealt with and the refresh re-run, and completeness was checked by hand connector by connector before we called it done.

The on-call engineer for the release was Kieran, in his seventh week at B-Brain. He raised the issue promptly and kept a clear log, which made this review much easier to write.

Cause

Root cause, not symptom: the incremental-load scheduler change shipped in release four-twelve altered how the scheduler decided what each connector should load. Under that logic the refresh could finish its run without loading everything it should have, and nothing in our monitoring noticed the difference between "ran" and "completed".

There are two contributing factors, and we would rather name them than hide behind the first.

  • No completeness check. We alert when the refresh job fails. We do not alert when the job succeeds but loads less than expected. A refresh that finishes quietly and incompletely looks identical to a good one.
  • Release cover. The release went out in the evening with a junior engineer on call and no senior platform engineer available to take the first call. Since the departure of a senior colleague in May, the platform team has had one senior engineer. That is a structural weakness, and it is on us.

What went well

  • The problem was identified the same working day it was visible.
  • Nobody said a fix was done until it had been checked against the data.
  • Customer success had accurate, plain information to take to customers.

What did not go well

  • We found out from the morning rather than from an alert.
  • The release was scheduled at a time when the on-call arrangement could not absorb a problem of this kind.
  • The release notes described the scheduler change accurately but we had not agreed with ourselves what "good" looked like on the first morning, so there was nothing to compare against.

Actions

ActionOwnerStatus
Refresh-completeness alerting: alert when the nightly refresh finishes but loaded less than expected for any connector, not only when it fails.Ewan PryceOpen
No release without a senior platform engineer on call. A release waits until one is named as cover.Ewan PryceOpen
Agree and write down, before each release that touches the refresh, what a healthy first morning looks like.Kieran Walsh, with Ewan PryceProposed

Both main actions were open at the time of this review. Neither is done until it has been shown working, and we will report on them in the same way we have reported on this incident.

Customer communication

Customer success are speaking to affected customers directly. The line they have is the one in this report: the brain was available, the data was stale, the cause was ours, and we are putting alerting and release cover in place. They have been asked not to promise more than that.

Not yet known

  • Whether any customer made a decision in the affected window on the basis of stale answers. We will only find that out by asking, and customer success are asking.
  • Whether other scheduler changes in the pipeline carry a similar risk. We are checking before anything else touching the refresh is released.

Sign-off

Ewan Pryce, Head of Platform Kieran Walsh, Platform Engineer

Reviewed Wednesday 29 July 2026

Ewan

Documentsmade from

Unusual terms

Current to 9 Oct 26
  • Freshness failure without an availability breach
  • Only Enterprise credited for freshness

B-Brain is a fictional company; every organisation and person here is invented. Built by site/build_site.py from the site tree, data as of Fri 9 Oct 2026. Help & Support

Help & Support

Open as a page

Help & Support

B-Brain is one place to read everything the company knows about its customers: the CRM, calls, emails, support tickets, product usage, invoices, documents, news and HR. Every page is built from those systems and the data is current to Fri 9 Oct 2026.

How to use the site

How to ask

Press Ask Brain in the header. Type a question, or pick one of the examples.

The site itself does not call an AI model; answers in Claude come from the same figures you see here.

What the data covers

DataRecords
Organisations75
People at customers290
B-Brain staff40
Calls784
Email threads1,169
Support tickets449
Documents477
Deals117
Invoices101
Events49
News items56

Data as of Fri 9 Oct 2026. Text marked AI was written by the brain from the records listed in its made-from link; an AI output that cannot cite its evidence is refused and the previous text kept. Where two systems disagree (for example a contract and the CRM), the key facts show both values and mark the difference.

B-Brain is a fictional company: every organisation, person and figure here is invented for this demonstration.

Who to contact

Email support@b-brain.example or talk to your B-Brain account team. Tell us the page address and what looked wrong; a screenshot helps.

Ask B

Ask B

B-Brain · read-only