PrettyModels AIAI Research Lab

Testing frontier AI against public-market outcomes.

We study one question: can large language models evaluate public companies with enough consistency to inform investment decisions? Our process converts the qualitative record of a business (filings, disclosures, market data) into structured scores, builds a portfolio from them, and measures the result in a public reference portfolio.

Marylin separates open-ended model judgment from deterministic portfolio construction. Around each reporting season, the same research sequence is applied across a defined universe: dated public evidence is assessed through a consistent analytical framework, converted into comparable scores, and translated by fixed portfolio rules into a reviewable monthly hypothesis. Inputs and outputs are retained for out-of-sample evaluation, while human review remains between the model portfolio and any real-world execution.

These figures are research observations, not investor returns or a performance promise. Comparisons may differ in fees, spreads, currency, taxes, risk and investability. Past performance is not a reliable indicator of future results. Read the Research & Risk Disclosures before interpreting them.