Trying the new voice assistant from AI startup Sesame is the first time I momentarily forgot I was talking to a bot.
Compared to ChatGPT's voice mode, Sesame's "conversational voice" feels natural, unforced, and engaging, which totally freaked me out.
On Feb. 27, Sesame launched a demo for its Conversational Speech Model (CSM), which aims to create more meaningful interactions with AI chatbots. "We are creating conversational partners that do not just process requests; they engage in genuine dialogue that builds confidence and trust over time," the announcement states. "In doing so, we hope to realize the untapped potential of voice as the ultimate interface for instruction and understanding."
Sesame's voice assistant is available as a free demo on the site and comes in two voices: Maya and Miles.
Since Sesame unleashed its voice assistant demo, users have reported awestruck reactions. "I've been into AI since I was a child, but this is the first time I've experienced something that made me definitively feel like we had arrived," user SOCSchamp wrote on Reddit.
"Sesame is about as close to indistinguishable from a human that I've ever experienced in a conversational AI," user Siciliano777 wrote on Reddit.
After talking to Sesame's bot, I was similarly wowed. I talked to the Maya voice for about 10 minutes about the ethics of using AI as a companion and came away feeling like I had a genuine conversation with a considerate, informed person. Maya's speech had a natural cadence, using interjections like "you know" and "hm," and even making tongue clicking and inhaling sounds.
The strongest impression I got from interacting with Maya was that she immediately asked questions, engaging me in the conversation. The bot started our conversation by asking how my Wednesday morning was going (note: it was indeed a Wednesday morning.) In contrast, ChatGPT voice mode waited for me to talk first, which isn't necessarily a good or bad thing, but it intrinsically shaped the conversation as me using ChatGPT as a tool for something I needed.
Maya asked about the risks of AI companions getting "too good at being human." When I told her I was concerned about the rise of more sophisticated scams and people losing touch with reality by replacing humans with bots, she responded thoughtfully and pragmatically. "Scammers are gonna scam, that's a given. And as for the human connection thing, maybe we need to learn how to be better companions, not replacements, you know, the kind of AI friends who actually make you want to go out and do stuff with real people," said Maya.
When I had a similar conversation with ChatGPT, I received a response that felt more like boilerplate language from a school guidance counselor: "That's a valid concern. It’s really important to balance technology with real human interactions. AI can be a helpful tool, but it shouldn't replace genuine human connections. It’s good that you're thinking about these issues."
While OpenAI pioneered voice mode's ability to be interrupted and have a more fluid back-and-forth conversation, ChatGPT still tends to respond in complete sentences and paragraph blocks, which sounds, well, robotic. When using ChatGPT voice mode, I never forget that I'm talking to a bot, and that's reflected in the conversation, which can feel stilted and forced.
By comparison, AI for Humanspodcast co-host Gavin Purcell posted a Sesame conversation on Reddit where it's practically impossible to distinguish which voice is the bot. Purcell prompted the Miles voice by telling it to act like an angry boss.
A very silly conversation followed about money laundering, bribery, and a mysterious incident in Malta. Miles didn't miss a step. There was no perceptible latency, and the bot remembered the context of the conversation and creatively advanced the improvisational argument by escalating, calling Purcell "delusional," and firing him.
Of course, there are some limitations. Maya's voice glitched a few times throughout our conversation, and it didn't always get the syntax right, like saying, "It's a heavy talk that come."
According to its technical paper, Sesame trained its CSM (based on Meta's Llama model) by combining the traditional two-step process of training text-to-speech models on semantic tokens and then acoustic tokens, decreasing latency. OpenAI similarly used this multimodal approach to training voice mode. However, it has never released a dedicated technical paper on voice mode's inner workings — it only discusses voice mode in the GPT-4o research.
Knowing this, it's surprising how much better Sesame's model is at conversational dialog. However, Sesame's launch is just a demo, so it merits further scrutiny when the full model comes out. According to the demo announcement, Sesame plans to open source its model "in the coming months" and expand to over 20 languages.
Copyright © 2023 Powered by
I compared Sesame to ChatGPT voice mode and I'm unnerved-燕尔新婚网
sitemap
文章
2
浏览
98415
获赞
7639
Facebook sued by news media outlet over 'Russia state
An online media company identified as “Russia state-controlled” on Facebook is now suingGoogle Assistant finally lets you book Uber, Lyft rides
It's taken a few years, but the Google Assistant is catching up to other digital assistants to get yControversial bill allowing authorities to shoot down private drones heads to the president’s desk
The U.S. government will soon have authority to shoot down private drones considered a threat.FollowCathay Pacific hit with data breach involving 9.4 million customers
Cathay Pacific is the latest airline to be battered by a major data breach.In a statement released oMom goes to the bathroom for 45 seconds and returns to find her toddler on a treadmill
If you've been around little kids for even a second, you know their greatest threat is often themselMusk focuses on Model 3 success in Tesla earnings call
Tesla CEO Elon Musk saved all the drama for his Twitter account during Wednesday's earnings call. ItApple's T2 chip makes third
Teardowns of Apple hardware have repeatedly revealed just how difficult it is to repair. Apple doesnFacebook bans far
Far-right group Proud Boys and its founder Gavin McInnes have been banned from Facebook and InstagraDyson introduces air purifier that destroys formaldehyde
Remember the terrible smell in ninth-grade biology when you dissected a frog? That's formaldehyde, aStudy uncovers clever way to get people to eat their veggies
Turns out if you call beets "dynamite" and sweet potatoes "zesty," people actually want to eat theirReddit partners with Patreon to offer up a special flair, put a focus on creator communities
Attention redditors and patrons, Reddit and Patreon are partnering up.Reddit, the #5 most visited weAmazon's holiday toy catalog is an evil/genius way to make parents spend money
As a kid, I loved flicking through holiday toy catalogs.While Toys "R" Us catalogs may be no longer,Trump's trip to London gets a cheeky 'baby blimp' ad from Sky News
Many people in the UK believe Donald Trump isbaby, so in honor of his upcoming state visit LondonersMichelle Obama's latest Instagram post gives new meaning to squad goals
Shoutout to former First Lady Michelle Obama for continuing to inspire us to keep up with our #HealtWild parenting advice from the first man to win a paternity leave suit
Welcome toSmall Humans, an ongoing series at Mashable that looks at how to take care of – and