On the core of our system are two language fashions, every fine-tuned to attain all monetary textual content on two dimensions. The primary measure is vagueness—is the corporate being particular with its language or hedging? The second measures complexity—is that this a real technical disclosure, or is dangerous information being buried in complexities? The 2-pronged system is purposeful; an organization’s challenges could be wrapped in vagueness or complexity, typically each. That’s the reason the rating of a single variable can not inform you sufficient.
What issues most is deviation from a benchmark. We benchmark each firm in opposition to its sector friends and in opposition to its personal submitting historical past, figuring out declining traits and sector outliers. Evaluating from absolute scores can current bias within the outcomes of our fashions. By measuring deviation from peer averages as an alternative, we offer a safeguard which cancels out the potential of any bias.
These labeled sentences populate a data graph connecting every firm to its trade friends, their filings, and its personal submitting historical past. This lets the system transfer past asking whether or not a danger issue part has modified in any respect, to a extra exact query: has the corporate’s disclosure on a particular danger matter shifted, relative to each its friends and its personal prior submitting? A light-weight mannequin makes the primary interpretive cross over the extracted sections, working alongside the vagueness and complexity scores. At roughly 97% decrease value per token than a frontier mannequin, it’s low-cost sufficient to run throughout the outlined universe. Solely the place that first cross identifies a real shift is the submitting escalated to a frontier mannequin for the deeper learn.


