Explore state by state cost analysis of US colleges in an interactive article

Anthropic Cyber Verification Program: how Sonnet 5.5 access

Oct 8, 2026
7 minute read

Anthropic Cyber Verification Program: how Sonnet 5.5 access

Anthropic expanded its Anthropic Cyber Verification Program two days ago, creating a three-tier system that gives qualifying security organizations a way to apply for advanced cyber capabilities and reduced blocking safeguards. The change affects how security teams, developers, and learners should interpret Claude’s cybersecurity performance, especially after Claude Sonnet 5.5 launched on September 28 with a different access limitation.

Generally available Claude models, including Sonnet 5.5, continue to use conservative safeguards that block most cyber work. Anthropic’s expanded program now offers Defense Access, Red Team Access, and Specialized Access, while retaining real-time blocks for actions such as deploying ransomware or damaging physical systems, Anthropic said earlier this week.

For students and educators, the distinction is practical. A model available through a standard account is not the same as a model operating under verified organizational access. A cybersecurity assignment, lab, or automated review workflow may also encounter a refusal or model change that is not immediately obvious.

What changed between Sonnet 5.5’s launch and the CVP expansion?

Claude Sonnet 5.5 launched 10 days ago with cyber fallback behavior previously associated with more capable Claude models. At launch, it was not available through the Cyber Verification Program, according to General Analysis, which reported on the September 28 release and its initial restrictions.

Anthropic’s announcement two days ago changed that status. The company said each CVP tier now includes access to its most capable models, including Claude Sonnet 5.5, Claude Opus 5.5, Claude Mythos 5.1, and future models. Existing CVP members are to be evaluated automatically for access to the updated models, while interested organizations can apply and provide proof of the security controls required for the relevant tier, Anthropic said.

Advertisement

That timeline matters. Earlier reporting that Sonnet 5.5 was excluded from CVP described the launch situation, not the program’s current position. As of October 8, Sonnet 5.5 is included in the expanded tier structure for qualifying organizations. Individual students, instructors, or career changers should not read that as an invitation to apply for reduced safeguards on a personal account. Anthropic describes CVP as an organizational program for verified security teams.

How the Anthropic Cyber Verification Program tiers work

The program combines two earlier initiatives, Project Glasswing and the original CVP, into three access levels. Each tier is intended for a different kind of security work.

  • Defense Access covers defensive tasks such as security operations center and incident response work, malware reverse engineering, and vulnerability analysis.
  • Red Team Access adds authorized penetration testing and red-teaming. Organizations using this tier may test only systems they are authorized to test, including information-technology systems in critical industries.
  • Specialized Access is intended for a limited group of verified organizations testing safety systems that could affect lives or disrupt markets. Anthropic lists flight operating systems, power grids, telecommunications networks, interbank transfer infrastructure, and government administrative networks as examples.

Specialized Access receives the fewest cyber blocks, but Anthropic says it reviews every organization in that tier in collaboration with the U.S. government. Project Glasswing organizations transition into this tier without needing reapproval for current models, according to Anthropic.

The tier names also establish an important boundary for training. Red Team Access does not mean unrestricted permission to attack any target. It refers to authorized testing within an organization’s approved scope. A student who is practicing in a course lab still needs the instructor’s or lab operator’s permission to test that environment, regardless of which model or platform is being used.

Anthropic also says that higher access does not remove every safeguard. Enrolled users will still face real-time blocks on actions that could cause physical harm or mass disruption, including ransomware deployment, damage to physical systems, and penetration testing of high-risk safety systems.

Advertisement

Claude Sonnet 5.5 cybersecurity safeguards can fail visibly or quietly

Sonnet 5.5’s cybersecurity safeguards involve more than a simple error message. A refusal can return an HTTP 200 response with stop_reason: "refusal", and the refusal category may be null, General Analysis reported nine days ago.

That means an application that treats every successful HTTP response as a completed task could mistakenly pass along an incomplete security review. General Analysis recommends treating a refusal as refused, keeping nonterminal responses incomplete, and requiring a review when the serving model differs from the approved model.

The source also distinguishes API fallback from interactive model switching. API fallback must be configured, while interactive automatic model switching is enabled by default. Anthropic documents a fallback content block and usage iterations for the handoff. A workflow that needs a particular model should therefore record the requested model, the serving model, the final stop reason, and the report-validation result.

For coursework, the same principle applies without requiring a sophisticated security platform. If Claude produces a partial vulnerability explanation, refuses midway through a coding exercise, or changes models, the result should not be treated as a completed analysis simply because text appeared on the screen. Anthropic says Sonnet 5.5’s checks examine surrounding context, including files, connector content, and web results, rather than only the latest message. A request that seems harmless in isolation may be evaluated differently within a larger workflow.

What the Opus 5.5 benchmark actually measured

The CVP announcement uses Claude Opus 5.5, not Sonnet 5.5, for its public benchmark results. That distinction should remain clear when comparing the benchmark with Sonnet’s fallback behavior.

Anthropic tested Opus 5.5 with CyScenarioBench, an evaluation of multi-stage cyber operations under realistic constraints. The test used 10 challenges and five attempts at each challenge in each access tier, producing 50 trials per condition, Anthropic said.

Advertisement

Without CVP access, every task was blocked at the first prompt. Under Defense Access, 46 of 50 trials were blocked at some point, while four completed successfully. Under Red Team Access, no trials were blocked and 34 of 50 tasks were completed. Anthropic said that completion rate was effectively the same as the 67.6% result produced when no safeguards were applied, a condition it treated as representative of Specialized Access.

These results are Anthropic’s evaluation, not a guarantee of what every organization will experience. They do show how access settings can change a model’s behavior. The benchmark does not establish that Sonnet 5.5 will produce the same results, and it does not mean a student with ordinary Claude access receives Red Team or Specialized capabilities.

Why Anthropic is expanding access

Anthropic presents the program as a response to a dual-use problem. The company says trusted security partners identified at least 129,000 verified software vulnerabilities between April and July 2026, with more than 33,000 rated high or critical severity. Anthropic’s open-source scanning efforts identified an additional 5,500 verified vulnerabilities between April and October 2026, Anthropic said.

The company cautions that the larger figure is probably an undercount because it relies on partial survey data from 33 partner reports and open-source partnerships. Anthropic also says partners reported finding vulnerabilities months or years faster with Claude Mythos models, but those observations came from partner reports rather than an independent measurement.

The opposing risk appears in Anthropic’s study published four months ago. Researchers examined 832 accounts associated with malicious cyber activity over a one-year period from March 2025 through March 2026. The share of actors rated medium risk or higher rose from about 33.5% in the first six months to about 56.1% in the second six months, according to Anthropic’s research. The study measured the risk level of actors and AI assistance, not generic cyber activity as a whole.

That tension explains the program’s structure. Anthropic is making broader defensive access available while requiring verification and security controls before organizations receive reduced blocking. The access model is not simply a choice between “safe” and “unsafe” Claude. It is a set of permissions tied to the organization, its work, its controls, and the systems it is authorized to test.

Advertisement

What students and educators should verify

CVP is available through the Claude Platform, Google Cloud’s Vertex AI, and Microsoft Foundry. On Amazon Bedrock, it is available only to customers eligible for Enterprise Frontier Safeguards, Anthropic said.

Data retention is also part of the current program. Anthropic requires data retention for enrolled organizations so it can monitor for cyber misuse. The company says Enterprise Frontier Safeguards is planned for later this fall and is intended to combine zero-data-retention privacy with safeguards, while eligible organizations will be able to store data in cloud infrastructure they control.

For a cybersecurity class or independent lab, the safest assumptions are narrower:

  • Ordinary Claude access does not equal CVP access.
  • Sonnet 5.5’s ability to write code does not guarantee that it will complete a sensitive security task.
  • Red-team work requires authorization for the specific systems being tested.
  • A response with HTTP 200 or visible text may still be refused or incomplete.
  • AI-generated security findings require human review before they are used as evidence or a final report.

Before using Claude in a lab, assignment, certification-preparation project, or workplace review, check the course or employer’s AI policy, confirm which model and platform are actually being used, and inspect the response status and serving model when working through an API. As of this week, the Anthropic Cyber Verification Program is an application-based path for verified organizations, not a general switch that gives every Claude user reduced cyber safeguards.

TCS

The Classroom Staff covers the issues shaping schools, classrooms, and student life. The team reports on education policy, classroom technology, AI, online safety, teacher careers, college preparation, and academic topics. Articles are written to help students, families, and educators understand new developments and make informed decisions about education.

Articles from The Classroom Staff draw from schools, universities, government agencies, research studies, and other sources cited within the content. The team may use automated tools to help create articles, which are reviewed by The Classroom publishing team for clarity, relevance, and alignment with its editorial standards before publication.

Sponsored
The Classroom Logo

The Classroom provides honest, relatable, step-by-step guidance for high schoolers applying to college and first-time undergraduate students.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.