Each week we find a new topic for our readers to learn about in our AI Education column.
We talk a lot about artificial intelligence in the financial services industry here at AI & Finance, just not in AI Education, where we’ve embraced the opportunity to write about the greater AI space. Well, this week we’re once again trying something a bit new and different because we’re going head-first into the financial services industry—specifically, what happens when you turn the technology behind generative AI onto financial data.
We’re going to talk about large transaction models, but first, let’s be clear—most of the applications to date of AI within financial services have been using language models—that is, AI technology that has been trained with large volumes of text to retrieve data and to speak and write like a human. These language models are powering recommendation engines and chatbots and AI agents being released across the industry, but they’re in essence little different from the reasoning and language models already deployed in other industries.
We’ve talked about large language models before—models like Claude and ChatGPT that power the popular AI chatbots most of us have interacted with at some point. We’ve also discussed some of the technology that underpins them like transformers and neural networks. Today we’re going to discuss what happens when you take the same AI concepts and technology, but instead of applying them to language like we would with a large language model, we apply them to the data being collected by financial institutions.
How We Got Here
Well, we wanted to write about a financial AI topic for once in AI Education, that’s really how we got here. In 2023, Cambridge, U.K.-based Featurespace launched TallierLTM, what it billed as the first Large Transaction Model. Since then, others have jumped into the space, including RBC via its Borealis division. Still others are taking a similar approach using large language models to understand and categorize transactions without training an actual large transaction model.
It turns out we’ve done a lot of the conceptual groundwork to discuss large transaction models, but we’re going to get a quick refresher. Large language models, as we’ve discussed, are based on transformers, a form of deep learning architecture that allows the model to use a system of artificial neural networks to recognize words, understand the relationships between words, phrases, sentences and paragraphs, and to find different meanings within a document or group of documents.
We train large language models by teaching them to accurately predict the next words in a sequence—the computer turns words and languages into math, but while traditional machine learning assigns a unique, unchangeable value to words, transformer technology is able to assign multifaceted values to words, enabling a computer to understand how words are related to one another, and to distinguish between words with similar sounds, spellings or meanings. While most computer systems are built to perform calculations efficiently, transformers are built to understand relationships and contexts and create nuanced, layered hierarchies. Thus, after an LLM has been trained on a sufficient volume of text, it is able to answer questions and write responses as if it were a human being.
So What Is an LTM?
As their name suggests, large transaction models are trained on transaction data, finding the patterns and relationships in the huge store of information kept by banks, asset managers and other financial institutions. Financial institutions are already doing much of this work the hard way—they’re retrieving and processing transaction data and teaching software to make decisions using more manual processes and workarounds. They’re looking at fragments transaction data like oldfangled machine learning looked at words—via an understanding of firm, fixed and unique values instead of flexible, multifaceted and mutable values.
Rather than relying on data that has likely been pulled from several silos, processed and sanitized, large transaction models (let’s call them LTMs) are able to analyze raw transaction data to find relationships. LTMs also learn as they go—which is important when your software may have to work across multiple geographies and answer to different regulatory requirements, or if it has to adjust to different types of fraud over a period of time. When it comes to fraud detection, the ability to learn over time allows for more nuanced understanding of anomalous financial behavior.
LTMs like Tallier work by evaluating transactions in the context of previous transactions to understand a sequence of financial behavior, then attempt to make a prediction on the characteristics of the next financial transaction based on past behavior, checking its own work and refining itself over time—not unlike a large language model trying to guess the next word in a sequence. It’s important to note that an LTM does this not just on the enterprise or business level, but also on the household, client and account level, giving it a granular perspective on normal and anomalous financial behavior. To prevent narrowing the model’s focus into the short-term, Tallier is also programmed to give more weight to longer term behavioral sequences.
What Do LTMs Do?
In TallierLTM’s case and elsewhere, the primary use for the large transaction model so far has been in fraud detection, where it is reported to have offered significant improvements in fraud value detection.
Hawk, another LTM provider, notes that in addition to anomaly and pattern detection, large transaction models also help banks and other institutions reduce the number of false positives—incorrectly flagged transactions—that need to be reviewed by human anti-money laundering, fraud and compliance teams. An LTM can also help financial institutions with more stringent know-your-customer requirements that may be on the horizon, accordiing to Hawk.
LTMs may have future applications in compliance, risk management and customer service, and may help financial institutions automate more workflows across different channels of business. For example, RBC is building an Asynchronous Temporal Model, or ATOM, as a foundation model for financial services, trained on transactions, that can be deployed to power a range of applications, including some of the tools within NOMI, its personal banking product, where the model is used to forecast upcoming debits and credits from accounts and understand cashflows, and to help clients input information.






