Physicists Neil Johnson and Frank Yingjie Huo at George Washington University have published a formula in the journal Patterns that estimates the number of correct responses an AI chatbot will generate before producing harmful content. The study builds on a preprint released in February, which demonstrated that the formula correctly predicted whether a model would tip immediately or after a delay in 15 of 16 clear-cut cases, representing a 94% success rate. The research focuses on the attention head mechanism, where accumulated context pulls the model toward specific answer clusters until it reaches a tipping point denoted as n*. This metric represents the count of good tokens produced before the first bad one appears.

The initial tests utilized six open-weight models from OpenAI, EleutherAI, and Meta, ranging from 124 million to 410 million parameters. The published paper expands this scope to seven models with up to 12 billion parameters, though these remain small by current industry standards. The primary application targets on-device AI systems, such as companion chatbots running offline on phones or laptops without cloud connectivity. To address safety gaps in offline environments, the authors propose a parallel monitor that flags when n* falls below a safety threshold. They also suggest methods to delay tipping points, such as injecting content into conversations, noting that while alignment training can shift tipping behavior for specific prompts, it cannot remove the underlying mechanism.