Wikiwand AI

Talk:2023 Capita data breach/GA1

From Wikipedia, the free encyclopedia

GA review

Article (edit | visual edit | history) · Article talk (edit | history) · Watch

Nominator: Joereddington (talk · contribs) 09:48, 13 February 2026 (UTC)

Reviewer: Metalicat (talk · contribs) 21:19, 11 March 2026 (UTC)


review by Metalicat

I'm reviewing this article against the good article criteria. The article covers a significant and well-documented incident, and the sourcing from the ICO penalty notice is strong. There are, however, several issues I'd like to see addressed before I can pass it.

Criteria notes:

  • 1a. Prose: Lead is overloaded; timeline is too granular for encyclopaedic prose; some duplication between sections.
  • 1b. MoS: Minor issues noted below.
  • 2. Verifiable: Well cited but leans very heavily on one primary source; a couple of weaker sources carry significant claims.
  • 3. Broad in coverage: Some important context and aftermath coverage is missing or underdeveloped.
  • 4. Neutral: No obvious concerns.
  • 5. Stable: No issues apparent.
  • 6. Illustrated: No images expected for this type of article.

Issues to address

1. Lead

The lead currently tries to carry the entire article. It includes the attack narrative, the ICO investigation findings, the 6.6 million figure, the £14 million fine, the High Court claim, and the number of claimants, all in a few sentences. The lead should summarise the article cleanly without front-loading every number. I'd suggest trimming it to: what happened, who was affected in broad terms, and what the main regulatory outcome was. The litigation detail and specific data volumes can live in the body.

The £14 million fine currently appears in the lead, in the impact section, and twice in the regulatory action section. That needs consolidating.

2. Timeline section

The minute-by-minute timeline reads like an incident log rather than encyclopaedic prose. The key events are: initial compromise on 22 March, a 58-hour delay before the alert was actioned, lateral movement over the following days, exfiltration of approximately 973 GB on 30 March, and ransomware deployment on 31 March. Those could be told in two or three prose paragraphs. Individual timestamps such as 07:52, 08:00, 12:21 are more detail than a general reader needs unless each one is important to understanding why the breach succeeded.

If the minute-level detail is retained, I'd want a clearer editorial reason for it. At present it feels like the ICO penalty notice has been transposed rather than summarised.

3. Source diversity

The article draws very heavily on the ICO penalty notice for the background, timeline, data volumes, and regulatory findings. That is a strong and appropriate primary source, but the article is close to becoming a paraphrase of a single enforcement document. I'd like to see a little more independent secondary sourcing woven into the body, especially for context, attribution, and aftermath, so that the article does not read too much like a close paraphrase of the ICO notice.

Two sources in the litigation section are a bit weak:

  • The Barings Law website is a claimant-side source. It is fine for Barings Law brought a claim on behalf of X people but should not be the sole basis for claims about the scale or significance of the litigation.
  • The Business Matters citation for the February 2026 High Court ruling is not the strongest source for a legal development. A law report, court listing, or broadsheet report would be better if available.

4. Coverage gaps

Three areas seem underdeveloped and are worth checking against the available secondary sourcing:

Capita's public sector role. The background section says Capita is a BPO provider but does not really explain why this breach mattered so much. Contemporary coverage notes that Capita handled major public-facing functions, including pension administration and services connected to the military and NHS, which helps explain why the incident affected so many people and drew major regulatory attention.

Attack attribution. Reliable secondary sources appear to attribute the attack to the Black Basta ransomware group, but the article body does not currently address this. If that attribution is included, it should of course be stated with clear sourcing and attribution in the prose.

Separate security exposures. There was also reporting in May 2023 about an exposed Capita Amazon S3 bucket. If that is outside the intended scope of this article, it may still be worth making clear that it was a separate incident, as readers could otherwise conflate it with the ransomware breach.

5. Inconsistent claimant numbers

The lead refers to more than 5,000 people in the High Court claim. The regulatory action section refers to more than 8,000 individuals. Both may be correct at different dates as the claim group grew, but as written this looks inconsistent and needs clarifying. If the number changed over time, that should be stated clearly, for example: The claim, initially brought on behalf of more than 5,000 individuals, had grown to more than 8,000 by February 2026.

6. Minor prose issues

  • the 23-24 year should be written out: the 2023–24 financial year.
  • Initially, 827.25 MB of data was downloaded; this eventually reaches 1.76 GB on this channel shifts from past to present tense.
  • all 59,359 accounts on the system is very precise. If that wording comes directly from the ICO notice, consider attributing it; otherwise round it or say all accounts on the affected domain.

Summary

This is a solid article on a notable incident with strong sourcing from the ICO penalty notice. The main work needed is to tighten the lead, condense the timeline into more encyclopaedic prose, add a little more independent secondary sourcing for context and aftermath, and resolve the duplication and inconsistencies noted above. I do not think these are fundamental problems, but they do need attention before the article is ready to pass.

Happy to re-review once these are addressed.

Metalicat (talk) 21:19, 11 March 2026 (UTC)

Fabulous! Thank you so much for your review! I know that a dry data breach article might be a little unattractive to many reviewers so I really appreciate the time.
  • On the minor prose issues - I've fixed the financial year and the tenses. I was a bit confused by the 59,359 comment - it does come from the ICO report and is attributed to page 20? (I'm amused because a GA reviewer on a different recent article suggested I wasn't being precise enough :) )
  • Claimment numbers - Good catch I have switched to 8,000 in the lead.
  • I've removed the extra mention of the 14 million in regulatory action. That was strange.
  • For comments 1 to 4 - I think they are useful and carefully stated, but I'd like to plead WP:GANOT for the purposes of this review and come back to them afterwards when other sources appear. FWIW - to talk about your sourcing point more generally... Yes, I think that things like the ICO report are PRIMARY, but I also think they are independent and WP:RS. The actual issue here (and I'm quietly slipping into gear on a longer rant) is that broadly, unless forced by law, companies release almost zero information about cyber-attacks other than "It happened" and some things that will later turn about to be media spin when it reaches court. In the majority of cases the law _is_ the ICO reports. If we look at 2015_TalkTalk_data_breach (a recent beneficiary of the Feb GA push), four people went thorough full court cases and were sentenced, but we know almost nothing of technical detail from those sources (partly because of the way the UK court system works - indeed, we don't seem to know which of the suspects actually performed the DDOS) and instead everything comes from the ICO. All the secondary sources are just quotes from the ICO. Indeed, I only started this article because the ICO released its report.
Now, I'm sympathetic to the argument that it would be better to view the ICO's report through a secondary source that included commentary and editorial judgement on an agency that, while independent, does have its own motives and bias. Particularly one that included some other major source (court transcripts and so on, often someone like Wired has quite a lot of other good sourcing) but there wasn't a great example in this case. Professionally, what I probably should do is write up such a paper myself at the same time I write the Wikipedia article because then I'd have a few more bits of publication to put in my promotion application, but who has the time? If there were a couple of good academic papers and maybe a book as well, then I'd definitely writing quite a different article, perhaps with ambitions for FA. *shrug* Does that make sense? For me the Black Basta stuff doesn't quite get over the line into supported.
Joe (talk) 10:28, 12 March 2026 (UTC)
Thanks Joe, and you're welcome. I find these kinds of articles more interesting than their reputation suggests. You should see what I write about if you need help sleeping!
Taking your points in turn:
59,359 figure. You're right, that is attributed with a page reference. I withdraw the comment. all 59,359 accounts on the system is very precise. If that wording comes directly from the ICO notice, consider attributing it; otherwise round it or say all accounts on the affected domain.
Minor prose fixes. Thank you for turning those around quickly.
WP:GANOT and sourcing. That's a fair point, and I think you're right. The ICO penalty notice is a reliable, independent, published enforcement document and in the cybersecurity space it is often the best source available for technical detail. You're correct that most secondary coverage of incidents like this amounts to journalists summarising the same ICO findings, so requiring secondary mediation of every claim would be an artificial hurdle rather than a genuine quality improvement. I'd still encourage broadening the sourcing over time, particularly if a good long-form treatment or academic paper appears, but I'm satisfied that the current sourcing meets the GA criteria and has been an interesting niche for me to learn.
Points 1 to 4 more broadly. I stand by the suggestions as improvements worth considering in due course, but I think your points on GANOT are fair here. None of those points amount to a failure against the criteria. The article is well written, verifiable, broad enough in its coverage, neutral, and stable.
I'm happy to pass this article. Well done on pulling together a clear account of a complex incident.
Metalicat (talk) 19:04, 12 March 2026 (UTC)

Related Articles

Timelines

Top Qs

Fact Checks