Startup DreamersStartup Dreamers
  • Home
  • Startup
  • Money & Finance
  • Starting a Business
    • Branding
    • Business Ideas
    • Business Models
    • Business Plans
    • Fundraising
  • Growing a Business
  • More
    • Innovation
    • Leadership
Trending

With One Million Displaced, Lebanon Turns to Digital Wallets for Aid

April 21, 2026

Nvidia’s Trillion Dollar Prediction Marks AI’s Inflection Point

April 21, 2026

“Uncanny Valley”: OpenAI and Musk Fight Again; DOJ Mishandles Voter Data; Artemis II Comes Home

April 20, 2026
Facebook Twitter Instagram
  • Newsletter
  • Submit Articles
  • Privacy
  • Advertise
  • Contact
Facebook Twitter Instagram
Startup DreamersStartup Dreamers
  • Home
  • Startup
  • Money & Finance
  • Starting a Business
    • Branding
    • Business Ideas
    • Business Models
    • Business Plans
    • Fundraising
  • Growing a Business
  • More
    • Innovation
    • Leadership
Subscribe for Alerts
Startup DreamersStartup Dreamers
Home » Meet The AI Agent With Multiple Personalities
Startup

Meet The AI Agent With Multiple Personalities

adminBy adminApril 22, 202510 ViewsNo Comments3 Mins Read
Facebook Twitter Pinterest LinkedIn Tumblr Email

In the coming years, agents are widely expected to take over more and more chores on behalf of humans, including using computers and smartphones. For now, though, they’re too error prone to be much use.

A new agent called S2, created by the startup Simular AI, combines frontier models with models specialized for using computers. The agent achieves state-of-the-art performance on tasks like using apps and manipulating files—and suggests that turning to different models in different situations may help agents advance.

“Computer-using agents are different from large language models and different from coding,” says Ang Li, cofounder and CEO of Simular. “It’s a different type of problem.”

In Simular’s approach, a powerful general-purpose AI model, like OpenAI’s GPT-4o or Anthropic’s Claude 3.7, is used to reason about how best to complete the task at hand—while smaller open source models step in for tasks like interpreting web pages.

Li, who was a researcher at Google DeepMind before founding Simular in 2023, explains that large language models excel at planning but aren’t as good at recognizing the elements of a graphical user interface.

S2 is designed to learn from experience with an external memory module that records actions and user feedback and uses those recordings to improve future actions.

On particularly complex tasks, S2 performs better than any other model on OSWorld, a benchmark that measures an agent’s ability to use a computer operating system.

For example, S2 can complete 34.5 percent of tasks that involve 50 steps, beating OpenAI’s Operator, which can complete 32 percent. Similarly, S2 scores 50 percent on AndroidWorld, a benchmark for smartphone-using agents, while the next best agent scores 46 percent.

Victor Zhong, a computer scientist at the University of Waterloo in Canada and one of the creators of OSWorld, believes that future big AI models may incorporate training data that helps them understand the visual world and make sense of graphical user interfaces.

“This will help agents navigate GUIs with much higher precision,” Zhong says. “I think in the meantime, before such fundamental breakthroughs, state-of-the-art systems will resemble Simular in that they combine multiple models to patch the limitations of single models.”

To prepare for this column, I used Simular to book flights and scour Amazon for deals, and it seemed better than some of the open source agents I tried last year, including AutoGen and vimGPT.

But even the smartest AI agents are, it seems, still troubled by edge cases and occasionally exhibit odd behavior. In one instance, when I asked S2 to help find contact information for the researchers behind OSWorld, the agent got stuck in a loop hopping between the project page and the login for OSWorld’s Discord.

OSWorld’s benchmarks show why agents remain more hype than reality for now. While humans can complete 72 percent of OSWorld tasks, agents are foiled 38 percent of the time on complex tasks. That said, when the benchmark was introduced in April 2024, the best agent could complete only 12 percent of the tasks.

Read the full article here

Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Related Articles

With One Million Displaced, Lebanon Turns to Digital Wallets for Aid

Startup April 21, 2026

“Uncanny Valley”: OpenAI and Musk Fight Again; DOJ Mishandles Voter Data; Artemis II Comes Home

Startup April 20, 2026

AI Agents Are Coming for Your Dating Life

Startup April 19, 2026

China Is Cracking Down on Scams. Just Not the Ones Hitting Americans

Startup April 18, 2026

The 70-Person AI Image Startup Taking on Silicon Valley’s Giants

Startup April 17, 2026

The US Army Is Building Its Own Chatbot for Combat

Startup April 16, 2026
Add A Comment

Leave A Reply Cancel Reply

Editors Picks

With One Million Displaced, Lebanon Turns to Digital Wallets for Aid

April 21, 2026

Nvidia’s Trillion Dollar Prediction Marks AI’s Inflection Point

April 21, 2026

“Uncanny Valley”: OpenAI and Musk Fight Again; DOJ Mishandles Voter Data; Artemis II Comes Home

April 20, 2026

AI Agents Are Coming for Your Dating Life

April 19, 2026

China Is Cracking Down on Scams. Just Not the Ones Hitting Americans

April 18, 2026

Latest Posts

How Arizona-Based Lectric eBikes Is Dominating The D2C Market

April 17, 2026

The US Army Is Building Its Own Chatbot for Combat

April 16, 2026

This Startup Wants You to Pay Up to Talk With AI Versions of Human Experts

April 15, 2026

Robotaxi Outage in China Leaves Passengers Stranded on Highways

April 13, 2026

Duolingo’s Luis von Ahn Wants to Delete the Blockchain

April 12, 2026
Advertisement
Demo

Startup Dreamers is your one-stop website for the latest news and updates about how to start a business, follow us now to get the news that matters to you.

Facebook Twitter Instagram Pinterest YouTube
Sections
  • Growing a Business
  • Innovation
  • Leadership
  • Money & Finance
  • Starting a Business
Trending Topics
  • Branding
  • Business Ideas
  • Business Models
  • Business Plans
  • Fundraising

Subscribe to Updates

Get the latest business and startup news and updates directly to your inbox.

© 2026 Startup Dreamers. All Rights Reserved.
  • Privacy Policy
  • Terms of use
  • Press Release
  • Advertise
  • Contact

Type above and press Enter to search. Press Esc to cancel.

GET $5000 NO CREDIT