Smartphones have dominated the digital ecosystem for the past decade and a half, serving as black holes of attention and our most intimate personal companions. Yet from their inception, phones were designed for "people staring at them"—their entire logic stops at the screen.
AI’s needs are precisely the opposite: it requires continuous perception of the physical world—seeing what you see, hearing what you hear, always present, rather than waking only when you unlock the screen.
When AI truly becomes a foundational capability, it will inevitably break out of the screen and seek its own form. This will be a long process of exploration and evolution.
The column “AI Artifact Chronicles” emerges from this context. iFanr aims to observe with you: how is AI transforming hardware design? How is it reshaping human-machine interaction? And more importantly—what form will AI take as it enters our daily lives?
This is the 19th article in the “AI Artifact Chronicles” series.
Over the past two years, the most noticeable change in AI isn’t just that models have become smarter, but that its positioning has shifted from a niche professional tool to a fundamental capability of everyday life.
Today, we’ve grown accustomed to issuing commands to AI anytime—while walking, driving, cooking, or even taking a break.

Image | Google
But as usage frequency increases, the traditional human-computer interaction model centered on keyboard and mouse from the PC era is gradually revealing its limitations:
“Typing manually” requires users to first form an idea, then organize scattered thoughts into coherent language, maintaining logical coherence throughout narration—only then can mental concepts be translated into prompts understandable by AI.

The Passion of Creation | Leonid Pasternak
On the other side of the keyboard, voice input during the smartphone era faces virtually no such constraints.
After all, the fundamental difference between speaking and typing is that speech allows people to think, supplement, and revise while outputting—making it far closer to the natural flow of thought than typing.

Beyond requiring stable desktop environments or sustained head-down posture, voice input also doesn’t demand strong writing skills—
For the vast majority of people without formal writing training, voice input is much closer to “truly natural interaction.”
Under such widespread demand, a wave of products centered around voice entry has naturally emerged.
For instance, Plaud, gaining momentum, packages recording, transcription, and summarization into a complete workflow;
While Typeless, becoming a reference product, removes filler words and performs format cleaning to “help you write.”

Wearable devices like Guangfan Lightwear go further, aiming to create a full-scene, end-to-end AI usage flow beyond traditional smartphones.

Image | Guangfan Tech
Though these products vary in form, they all bet on the same future:
The next generation of AI entry points won’t require users to sit down, open an app, and type carefully.
Natural language interaction sounds highly advanced, but its real advantage lies in communication efficiency—not groundbreaking technology.
The lowest-barrier solution is simply a pair of TWS earbuds plus an AI on your phone, forming a 24/7 virtual assistant.

Original AirPods ad “Bounce” | Apple
After all, earbuds already reside in your ears, and mobile operating systems and third-party apps already offer built-in speech-to-text, shortcut commands, and cross-app input.
No dedicated hardware required—you can relay fragmented commands to any AI model instantly.
Even more harshly, its speed, compatibility, and cross-device capabilities still outperform most so-called “revolutionary” dedicated AI hardware on the market:

Image | Mashable
Yet reality reveals that this nearly barrier-free combination’s true bottleneck isn’t the phone—but the microphone’s handling of ambient human voices.
Taking AirPods as an example, due to their emphasis on natural sound quality and reluctance to apply aggressive noise suppression, they often capture surrounding conversations along with your speech, leading to inaccurate transcriptions.
This may not matter much for calls or voice messages, but for AI input, if the prompt is “undercooked,” the response is likely completely irrelevant:

In contrast, domestic TWS earbuds generally adopt more aggressive voice isolation strategies, sacrificing some transparency mode naturalness for better accuracy during phone calls and AI voice interactions.
Looking deeper, this implies a shift in how we evaluate earbuds: as AI becomes the dominant use case, the criteria for judging audio quality are changing.
Previously, earbuds were judged by sound quality, latency, and noise cancellation. In the future, besides these, we must also compare the strength, accuracy, and separation capability of voice noise reduction.
We might even say that earbud microphones are evolving from forgotten call accessories into frontline AI controllers.

Mission: Impossible (1996)
The second path for voice control lies in smart glasses.
Despite constant claims of adding a HUD to life, commercially successful smart glasses remain fundamentally basic models centered on voice input, photography, and open audio.

Image | Laptop Mag
The reason is straightforward: display systems in glasses sacrifice weight, battery life, and cost, and require cultivating entirely new user habits from scratch.
Meanwhile, “see something, ask AI, get an answer”—this alone constitutes a clear, compelling AI value proposition.

Image | Meta
But the issue with smart glasses is that their “effortless” nature exists mostly in promotional videos.
After all, when a camera-equipped pair of glasses constantly faces others, it’s hard for people to tell whether the device is recording or capturing audio—this uncertainty itself creates social discomfort.
Additionally, constrained by size and battery life, the microphone arrays, call noise cancellation, and speaker performance of smart glasses aren’t necessarily superior to mature TWS earbuds.

Image | Stuff
More critically, under current sales models, these products are often deeply tied to brand clients and cloud-based models, leaving users little freedom in model selection, data migration, or cross-device switching.
The simplest example is Meta.
Meta Ray-Ban smart glasses are well-balanced across features but force users to connect to Meta AI, unable to directly trigger Siri or other assistants via Bluetooth as a standard audio device.

Image | Engadget
This level of binding means that while smart glasses represent the archetype of voice-based AI interaction, they struggle to become mainstream products.
The third path—gaining traction since 2026—is creating dedicated hardware entry points for voice input to AI.
For example, OpenAI’s recent collaboration with Work Louder on Codex Micro is an external control console designed specifically for AI agents.
It includes a dedicated push-to-talk (PTT) button for voice input:

Image | OpenAI
Although Codex Micro lacks a built-in microphone and relies on other audio devices connected to a computer for PTT input, this design carries symbolic significance:
Voice input is no longer just a small icon in the UI—it is gradually acquiring a stable physical position akin to a right-click mouse, a phone camera button, or a fingerprint sensor on a keyboard.
Setting voice input as a dedicated button essentially acknowledges its high-frequency nature—possibly even more frequent than the 26 letter keys for certain users.

Luo Yonghao demonstrating TNT
Going further, the vision behind PTT/TNT is to completely eliminate screens.
iFanr previously reported that among several hardware projects OpenAI is planning, one includes a screenless portable device that understands context through voice, camera, and environmental sensors.

Image | MacRumors
If OpenAI truly designs this way, it would effectively validate the path taken by Humane AI Pin.
Certainly, backed by ChatGPT and Codex, OpenAI’s badge would be far more practical than Humane’s.
Meanwhile, researchers are developing innovative solutions to address two persistent challenges in AI voice interaction: the awkwardness of stating personal plans in public, and recognition errors in noisy environments.
For instance, recent “ultrasound tongue motion recognition” devices use ultrasound probes placed against the jaw to detect tongue movements, then decode them into language using machine learning:

Image | Hackaday
Other studies attempt to embed sensors into glasses or headsets, detecting subtle facial, cheekbone, and mouth movements to recognize silent commands—Apple quietly acquired Q.ai, which was researching this direction.
Related reading: Apple’s second-largest acquisition ever, whose target wasn’t a phone | Hard Philosophy
These devices remain far from consumer-grade products, with notable limitations in recognition range, accuracy, and wearing comfort—but the direction is already clear.
In the future, “voice input” may not actually require vocalizing at all—what we call voice interaction could ultimately evolve into direct recognition of human articulation gestures, lip movements, or facial muscle activity.

Image | Game Anim
Ultimately, why voice input gains precedence over keyboards in the AI era comes down to one fact: humans are not purely “visual animals.”
In other words, while we’re accustomed to receiving information visually, we’re not skilled at sending information via visual signals.
When transmitting information, we still follow a 300,000-year-old habit: making “coo-coo-ga-ga” sounds is more efficient than gesturing.

My Fair Princess (1994)
Using voice to command AI may seem inefficient—people often repeat themselves, stray off-topic, and expressions rarely achieve conciseness.
But its greatest strength is eliminating the sense of AI as a tool.
Whether users subconsciously treat AI as a subordinate, secretary, or companion, dialogue as a medium remains closer to humanity’s innate collaborative nature than filling out command boxes:

Yet while “voice interaction” is convenient, this convenience comes with rising hidden costs.
Traditional STT tools only convert sound into text; the final model charges based solely on the resulting text input and output tokens.
Tools like Typeless, however, invoke additional models to clean up and rephrase spoken language before sending the refined content to another AI, adding another layer of token consumption.

Image | MakeUseOf
With multimodal models now entering “native audio” stages, processing continuous audio, context, tone, interruptions, and outputs causes token consumption per task to be ten times or more that of pure text.
Although compared to high-definition image and video processing, the total token cost for audio remains manageable, it is no longer a negligible feature.
Especially in prolonged real-time conversations, users can’t easily control message length like with text input, yet models must continuously maintain connection, understand context, and generate responses.

Image | Mashable
This leads to a paradox: the more natural AI voice interaction feels, the more complex the underlying computation becomes, resulting in higher user costs.
This almost constitutes a kind of “open conspiracy” among recent AI companies:
Encouraging voice interaction brings multiple benefits, but simultaneously pushes token consumption to a new height—no vendor would oppose charging for it.
Thus, for model providers and accessory brands, future competition in AI services isn’t just about understanding human language.
More crucially, who can best balance massive cost disparities across modalities with lower latency, less redundant computation, and transparent billing?
After all, getting people to speak to AI is easy. The real challenge is making them want to keep speaking.
This article comes from WeChat Official Account “iFANR” (ID: ifanr), author: Ma Fuyao
Source: iFANR
Disclaimer: Contains third-party opinions, does not constitute financial advice
a16z: 80% of AI Budget is Idle
4 days ago
OpenAI Former CTO Unveils 975B Open-Source Champion
5 days ago
DeepSeek Races for Sci-Tech Innovation Board Listing: Valuation Reaches $71 Billion
6 days ago
26 People Sue Meta: Will AI Fire You for Taking Maternity or Sick Leave?
6 days ago
Liáng Wénfēng Becomes the New Richest Man in AI, 36 People Get Rich Overnight from Large Models
7 days ago
Today's Leaderboard for Large Model Token Usage
05-08
Doubao Launches Paid Model: Can AI Subscriptions Work?
05-06






