TLDRocket
Sign in

Morgan Stanley is shaping the future of financial services

OpenAI

Morgan Stanley is leaning on OpenAI-style evals to grade its AI tools before advisors ever touch them. Turns out testing accuracy matters more than flashy demos when it's someone's retirement money.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Morgan Stanley has spent the past couple of years building AI tools for its financial advisors, most notably an assistant trained on the firm's own research and internal knowledge base. The harder problem was never getting a model to sound confident. It was proving the thing wouldn't quietly get something wrong in front of a client.

That's where evals come in. Rather than shipping a model and hoping, Morgan Stanley built a structured evaluation process to test outputs against real advisor questions, checking for accuracy, relevance, and whether the answer actually reflects the firm's compliance standards. Humans with deep subject knowledge review the results, flag gaps, and feed that back into how the tool gets tuned. It's less glamorous than a product launch, more like quality control on a factory floor, except the product is financial guidance touching billions of dollars in client assets.

The firm treats this as an ongoing discipline, not a one-time checkbox before launch. As the assistant expands into new tasks, summarizing earnings calls, drafting research recaps, prepping advisors for meetings, each new capability gets its own round of scrutiny. Wealth management is a business built on trust accumulated over decades, and a single bad answer repeated across thousands of client conversations could do real damage fast.

There's also a talent angle here. Morgan Stanley has framed evals as a way to bottle the judgment of its most experienced advisors and researchers, encoding what good looks like so newer employees get faster feedback loops. That's a quieter, more interesting use of AI than the usual chatbot pitch: not replacing expertise, but compressing the time it takes to develop it.

What stands out is how unsexy the actual work is. No one is talking about AGI or agents running trading desks. It's spreadsheets of test cases, disagreement logs between reviewers, and slow iteration. In an industry where regulators and compliance officers outnumber enthusiasm, that boring rigor might be exactly what determines whether generative AI sticks around in finance or gets quietly shelved after the first embarrassing headline.

My take — AI-written commentary, not fact-checked reporting

This is the part of enterprise AI nobody puts in a keynote, and it's the part that actually decides whether any of this survives contact with a regulator or an angry client. I'll take a boring eval spreadsheet over a slick demo any day, because in finance the cost of a hallucination isn't a meme, it's a lawsuit. If more firms borrowed Morgan Stanley's patience instead of chasing the fastest possible rollout, we'd hear a lot fewer AI horror stories.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.