2026-07-29 20:01
Introduction: Scarcity of compute power has spurred engineering optimization—K3 enables China’s open-source models to move beyond low cost and reach SOTA thresholds, sending shockwaves through Silicon Valley
Kimi’s K3 can be seen as the second major surprise from China’s open-source large models after the “DeepSeek Moment” during last Lunar New Year.
In several benchmark tests, K3 approaches or even surpasses some of the most advanced proprietary models from Silicon Valley. This has led to its nickname: “The Kimi Moment.”

But what we want to discuss in this article is not just Moonshot.
Over the past decade, China’s AI industry has experienced two generations of “Four Little Dragons”: the first generation, represented by computer vision companies like SenseTime, Megvii, Yitu, and CloudWalk; the second generation being Zhipu, Moonshot, MiniMax, and DeepSeek. Interestingly, these two generations are connected—through talent, capital, lab experience, and even trauma left from previous entrepreneurial cycles—all influencing today’s large model competition.
Our guest Kiwi once led the A1 round investment in Moonshot at Meituan Dragon Pearl, and has been a witness and investor throughout the evolution of China’s two AI waves. In this interview, she recounts her early investment story in Moonshot and retraces the history of China’s large model entrepreneurship with us.

Chen Xi: Hello, Kiwi. Welcome to Silicon Valley 101. Could you introduce yourself first?
Kiwi: Hi everyone, I’m Kiwi—a seasoned “jack-of-all-trades” in the AI industry for ten years. If I had to sum myself up in one sentence, I’d say I’m someone who continuously holds biases, keeps making mistakes, gets repeatedly proven wrong—but still stubbornly maintains those biases. I’m just an ordinary explorer in the AI field.

Chen Xi: You’re being too modest. I’m really looking forward to this interview because you’ve witnessed the full cycle—from the so-called “Four Little Dragons” of the last wave to the current “Four Little Dragons” or even “Six Little Tigers”—over roughly a decade. I’m eager to hear your insights on what happened in between, the people behind it, and the shifts in capital.
Kiwi: Sure, but let me emphasize—I’m not deeply technical or a model expert. I’ve just been incredibly lucky, always surrounded by geniuses, watching how brilliant minds explore and create the future.
Chen Xi: Your name actually fits Kimi perfectly. I only realized this while drafting the interview outline—“Kiwi” spelled backward is “Kimi.” Since many Silicon Valley model companies love naming their products after fruits or using internal codenames during testing, maybe next time they could consider using “Kiwi.”
Kiwi: Maybe it’s fate—or just coincidence. Actually, my Chinese name is Qi Yi, and back when I was overseas, I was often asked if I was from New Zealand.
Chen Xi: Let’s start with your investment story in Moonshot (Moonshot). K3 has recently attracted massive attention. You were at Meituan Dragon Pearl and invested in Moonshot’s A1 round—could you tell us the story?
Kiwi: This story begins in 2022, right after ChatGPT’s release.
After experiencing ChatGPT for about a week, we realized something was different. We initially thought GPT’s decoder-only architecture was just a niche branch of generative NLP tasks. But after prolonged interaction with ChatGPT, we sensed it might be a nascent form of AGI capable of passing the Turing test. That’s when I started discussing it with friends, asking for their perspectives.
Over the next few months, I spoke with every friend from the previous generation—researchers, engineers, and outstanding researchers emerging in NLP and reinforcement learning. After these conversations, we concluded this was a monumental shift. It had generalization capabilities unlike anything from the previous AI era. So we began tracking: if China were to pursue this, what kind of team would lead it?
To be honest, we felt a new playing field was needed. Legacy companies burdened with historical constraints couldn’t freely pursue something whose return on investment was unknown. The only anchor remained talent—those who truly believed in AGI and were genuinely excited by it would naturally gather. So our approach was simple: identify who believed in AGI and whose passion for technology and belief in AGI were strong enough.

Based on that, we gradually mapped out who in China was interested. At the top of our list was Kimi—Yang Zhilin. Honestly, I didn’t know him personally at first. I knew more people from the computer vision world, while Zhilin worked in NLP. But before ChatGPT’s release, if you looked at who was pursuing scaled models under the Transformer architecture in China, Zhilin stood out. For example, in 2019, during his internship at Google, he co-authored Transformer-XL and XLNet—names that already hinted at his focus on scaling.
Later, we noticed institutions like Beijing Academy of Artificial Intelligence, where models such as CPM and GLM emerged. We saw Tsinghua’s Professor Tang Jie, Liu Zhiyuan, and Zhilin all experimenting with decoder-only scaling. A year later, when Huawei launched Pangu 1.0, Zhilin’s team was still involved. You see, even before ChatGPT’s release, this person seemed consistently committed to this mission. So we were very eager to connect—but honestly, we tried for four months without success. Only after the first round closed did we finally manage to add Zhilin on WeChat, and after that conversation, we decided to invest.
Chen Xi: I heard getting added on WeChat wasn’t easy.
Kiwi: Part of it was my own shortcomings. I was truly obscure in the industry—just transitioned from AI industry to investing, with no prior deals under my belt. In that context, not being accepted on WeChat was entirely reasonable.

But I visited his former AI company three times. I met his other co-founders. Later, I chatted with Tim (Zhou Xinyu), who came from Megvii—part of the first generation of “Four Little Dragons.” Our conversations were more frequent. Tim told me he was joining Kimi (Yang Zhilin) to start a company. We discussed extensively. Kimi is Yang Zhilin’s English name, coincidentally sharing the same name as his model. But internally, to avoid semantic confusion, Zhilin now goes by KK.
I have many excellent friends who deeply believe in this LLM and AGI wave—they naturally gravitated toward each other. Some of my close friends had already joined Moonshot. I asked them: “What do you need?” They said: “We desperately need talent.” The initial team had more algorithm researchers, but large model development also required strong foundational systems and infrastructure (infra).
So I scoured the first generation of “Four Little Dragons” for colleagues I admired. To be honest, that cohort endured a long cycle—high highs and deep lows. Friends scattered across various places: some stayed with the “Four Little Dragons,” others went to big tech, others to quant firms. We reached out to each one, asking if they were interested in this new wave. Truthfully, 80% declined.
Having lived through a full cycle, having seen idealism shattered, they had already poured their most passionate years—late twenties to thirties—into AI once. Yet, from the 1.0 era, they hadn’t achieved a conventional success like a groundbreaking internet-scale company. So many developed a failure path dependency: perhaps they’d pause exploration or just observe for now.
But a small fraction—10% to 20%—didn’t care about personal gain. They were simply driven by excitement and the desire to contribute to something meaningful. These were the ones we persuaded to join Moonshot.
Later, I introduced several exceptionally talented engineers. These engineers demonstrated real value and capability during model training. They then vouched for us to Kimi (Yang Zhilin), saying we’d spoken with Kiwi, who was genuinely interested in the team.
Eventually, one day while waiting at a train station in Shanghai en route to Hangzhou, Yu Tao (CTO of Moonshot) and his friends told me Kimi agreed to meet us. They set up a group chat and added us on WeChat. We flew to Beijing the following week. There was definitely a rocky period.
Chen Xi: When was your first meeting with Yang Zhilin? What was your first impression?
Kiwi: Probably late April or early May 2023. My first impression was that he was a down-to-earth researcher. Many great researchers are like that—simple, pure, focused intensely on solving problems. The people I know, whether at Moonshot, DeepSeek, or Zhipu, include a core group who are genuinely passionate about the field. They are authentic.

I didn’t romanticize it. I vividly remember October 2023, when they released Kimi Chat and their first model. Because of exceptional long-context handling and strong alignment, I considered it one of the best-performing models in China at the time. After using it, I felt genuinely excited. I asked Kimi’s friends: “How did you achieve such results?” He replied: “We didn’t do much—we just diligently did everything that needed to be done.”
So my first impression of Zhilin was that he was purely talking about technology. When we asked about his understanding of large models and AGI, he shared that his 2023 goals boiled down to LTV: L for long context; T for truthfulness (reliability); V for multimodality, specifically video. He believed these directions were crucial for future AGI. Long context’s importance needs no elaboration; truthfulness addresses hallucination issues; and V reflects his view that once text modality is exhausted, we’ll need more data skills to augment intelligence. At that moment, I was simply learning from him—learning his perspective on AGI and large models.
Chen Xi: The LTV concept you mentioned is fascinating. In advertising, LTV stands for Lifetime Value—the total value a customer brings over their lifetime—an extremely commercial term. Yet for Zhilin, LTV is profoundly AGI-oriented. Same acronym, different meanings—quite intriguing. So when Zhilin finally agreed to meet us, did he become open and ready to discuss funding?
Kiwi: I don’t think he was closed before. The market simply has an extremely low signal-to-noise ratio. Someone like me, with no clear label, likely appeared as noise rather than signal. So perhaps our persistence and sincerity eventually caught his attention. This was just a touchpoint—this wasn’t the moment he suddenly opened up to fundraise.
Chen Xi: So he heard that there was an investor named Kiwi who kept persistently trying to talk to him, and who also referred many colleagues to the company, all of whom spoke highly of her.
Kiwi: Yes. I believe it was those friends who helped build trust. Ultimately, it’s about human trust. My label alone meant nothing—just an ordinary person. But because my friends had established trust, their endorsements gave us this small opportunity.
Chen Xi: So you added him three times?
Kiwi: I don’t recall exactly how many times. Maybe weekly, just another request among many.

Chen Xi: Let’s talk about your investment. Why did Meituan lead the A1 round?
Kiwi: Honestly, by the time we led, it was already May–June 2023. At that stage, other major model companies in the market had already raised their first to third rounds. Moonshot’s valuation was relatively low compared to peers—ZeroOne, Baichuan, which had pre-money valuations nearing $1 billion. Even earlier-established companies like MiniMax and Zhipu, having trained decoder-only models earlier, were valued above $1.5 billion.

Perhaps because Moonshot’s team was relatively young and hadn’t yet built a major company, the market paid less attention. So we were fortunate to get the lead role amid lower visibility.
Chen Xi: Was the internal investment committee discussion smooth?
Kiwi: From some angles, yes—it was fast after our internal IC (Investment Committee) discussion. But from others, it was extremely intense. This wasn’t just an internal debate at Meituan Dragon Pearl—it reflected broader skepticism across the market toward AI.
We believed a generalizable AI model would inevitably spawn a super application, even if we didn’t know what it looked like. The key question became: Who could build this super model plus super application? This was the most debated point. I recall Xing Ge (Meituan CEO Wang Xing) saying, “The CapEx here is enormous. If you believe in scaling laws, this is an infinite sink.”
For startups, this was completely beyond traditional venture capital capacity. We wondered: Would this become a big tech opportunity? Would ByteDance absorb all possibilities? Honestly, we didn’t reach consensus or find answers.
But I remember Xing Ge, on his way to another meeting, said: “I truly believe the probability of a startup creating a super model and super application is low, but I’m willing to support it anyway.” That single line pushed our decision forward. So, did we know a specific company would succeed? No—we didn’t. But we were willing to believe in idealism, which might still hold exploratory potential.
Chen Xi: Among so many model companies—MiniMax, Zhipu, etc.—why did Wang Xing and you bet on Yang Zhilin?
Kiwi: Returning to our criteria at Meituan Dragon Pearl: we sought companies that were genuinely pure in their pursuit of AGI and scaling.
Some companies’ goals included revenue, B2B, DAU growth. While these could align with model excellence at certain stages, they sometimes subtly conflicted. I don’t judge this—it’s just that the truly pure friends I knew were at Kimi.

Chen Xi: After Meituan’s lead investment, Kimi faced a tough year in 2024. We saw it begin investing in user acquisition, sparking polarized opinions. Why do I feel this trend? Because when we upload videos to Bilibili, our AI-focused content often features banners or ads from Kimi. At that time, revenue seemed critical—there was internal pressure to prioritize it. From an investor’s perspective, does this contradict their original promise of AGI?
Kiwi: I think it’s part of the exploration process. First, in early 2024, I recall a KOC (Key Opinion Consumer) pitching Kimi’s long-context research report reading ability to analysts. This instantly made non-AI audiences realize Kimi’s model was excellent, triggering a massive DAU surge. As DAU grew, new users arrived—bringing diverse tasks. These tasks, during reinforcement learning training, may have enriched the model with additional data.
Second, let me highlight: Kimi is an exceptionally difficult team. Or, for any non-big-tech-funded startup, pursuing AGI is inherently hard. AGI is expensive—this is a fact. To spend money, you must convince resource-rich parties to give you funds. How do you persuade outsiders to allocate resources?
Moreover, in 2024, Chinese startups faced extreme compute scarcity—this is a unique challenge for Chinese large model companies. But ironically, due to this scarcity, companies like DeepSeek, Kimi, and Zhipu perform extensive optimizations—sometimes achieving better performance than North American counterparts with abundant resources who don’t prioritize these optimizations.
Chen Xi: I think a harder challenge for Chinese model companies is the fierce ToC product competition. And most are free—especially after Douniu and Qwen launched. Now, it’s big tech supporting them. If Kimi pursues ToC, it directly competes with ByteDance, Alibaba, and even WeChat, which has recently shown activity. So the pressure must have been immense.
Kiwi: Absolutely immense. Worse, Kimi couldn’t charge $20 like ChatGPT—because it wasn’t a SOTA model. China hadn’t yet achieved SOTA. Both inference and training costs were high. With the fundraising scale of all model vendors in 2023, it was far from sufficient to pursue AGI.
So all startups had to balance: pursuing AGI versus proving to resource providers they were actively striving for AGI. These two goals aren’t always aligned. Every company faced this dilemma.
Chen Xi: When was the turning point for the user acquisition trend at Kimi?
Kiwi: I’d say early 2025 after R1’s release. Kimi’s first DAU spike wasn’t from ad spending—it came from the quality of its first model and product. Its long context and alignment were excellent, and users felt it immediately.
The model was strong. The product was strong. R1 also revealed the model’s reasoning process—something OpenAI hadn’t opened publicly. For the first time, the tech community and general users saw how the model thinks. This attracted massive interest. The open-source spirit here was truly remarkable.
So this was essentially a market education process—not for model companies, but for resource holders further from AI. The core message: build your model and product solidly. After R1, it taught resource providers: stop pressuring model makers with excessive DAU demands or immediate profitability. Instead, believe in AGI. At that stage, model companies aiming for AGI could refine their goals.
Chen Xi: Why didn’t Meituan participate in Moonshot’s 2024 round?
Kiwi: We observed—entirely due to severe GPU scarcity in 2024. Domestic compute resources were extremely limited. By late 2023 and early 2024, leading US frontier model companies had clusters with tens of thousands of GPUs. In China, only a few companies—DeepSeek, Qwen, ByteDance—had single clusters exceeding 10,000 GPUs. Even with ambition, you need resources. So our compute infrastructure wasn’t fully ready—this was likely why we hesitated.

Chen Xi: Is Kimi still facing compute shortages? Look at their recent client limits.
Kiwi: All AGI companies face compute shortages—OpenAI does too. Once, we asked a top-tier model vendor: “Do you lack compute?” They replied: “Compute is always fully utilized—give us more, we’ll use it all.”

Chen Xi: Compared to 2024, what significant progress has Moonshot made in 2025?
Kiwi: Overall, I think things changed significantly after o1. In 2024, scaling mainly relied on existing, recorded data—pre-training and SFT datasets. But after o1 in 2025, the focus shifted to reinforcement learning: define problems and rewards upfront, then use test-time scaling (TTS) to compute expansion, driving the next leap in intelligence.
What excites me most about 2025’s scaling is that the paradigm is shifting from relying solely on static, recorded data to allowing self-defined aesthetics to drive model advancement. With only static data, it becomes a resource race—scaling efficiency and quality matter. But when you can leverage your own aesthetic judgment to propel models forward, taste (taste) becomes paramount. So for vendors with limited resources and application-layer companies, a new wave of opportunities emerged. 2025 was truly exhilarating.
Chen Xi: It feels like everyone has returned to research.
Kiwi: Exactly. We’re redefining what constitutes a good task, a good problem, and how to evaluate them. Taste is multidimensional. At any given moment, there’s no absolute standard answer—so diversity increases. This is a particularly fascinating phase.
Chen Xi: But when K2 launched, the market response was lukewarm—now it’s “wow.” What was the team’s atmosphere like then?
Kiwi: I think we were already thrilled. K2 marked another milestone—China’s models crossed another threshold into a new phase.
Chen Xi: Wasn’t it the first to emphasize agentic models?
Kiwi: Actually, Claude 3.7 already emphasized agentic behavior. But K2 made Chinese models competitive with frontier models—with only half a generation or a few months behind—and offered exceptional cost-performance. That’s why K2 generated such excitement.

Chen Xi: Entering early 2026, we saw OpenClaw emerge—the “crayfish fever” in China sparked a surge in open-source models. Kimi’s API usage on OpenRouter also rose slightly.
Kiwi: Yes. During 2023–2024, people focused almost exclusively on frontier models, as their capabilities were steadily improving. Better models solved problems better.
But crayfish models are token-hungry and expensive. Most “shrimp farmers” found their tasks weren’t deeply complex—often not requiring solutions to the hardest scientific problems. So cost-effectiveness became a concern. At that time, K2.5 was an especially compelling choice.
Not just K2.5—many domestic models offered great value. Zhipu too. But recently, it’s no longer just about cost-effectiveness. What excites me about K3 is that prior to this wave, Chinese open-source models were often seen as “substitutes” due to their high cost-efficiency and superior inference optimization.

But K3’s launch is different. Consider K3’s training process—it surpassed the Opus series. Though Opus 5 has since been released (as of this interview date, July 25, 2026), at launch K3 already exceeded all models in the 4.8 series. Models truly comparable to K3—or even slightly better—were only recently released, like Fable. Thus, during its training, K3 surpassed all contemporaneous SOTA models. This time, Chinese models aren’t competing on cost-efficiency—they’re reaching SOTA. This gives me personal encouragement.
Chen Xi: And K3 is open-source. So this year’s “Kimi Moment” feels like a similar shock to Silicon Valley as last year’s “DeepSeek Moment,” but the urgency is greater.

Chen Xi: One interesting thing about Moonshot: I visited their Beijing office recently. The company feels cool—inside is a white piano with a Pink Floyd record. Every conference room bears the name of a rock band. Several founders founded rock bands at Tsinghua. Does this relate to AI at all?

Kiwi: I think at the core, if a company genuinely pursues AGI, it must be driven by genuine excitement and curiosity. Or, “pursuing something potentially useless” might reflect the essence of pursuing AGI.
Chen Xi: Not useful—that means no immediate product.
Kiwi: Even beyond products. For instance, when we started companies based on GPT-2 and GPT-3, we first asked: “What problems can it solve?” Then we designed the model—product-market fit first. But scaling itself—no one knows what emergent abilities will arise. So curiosity, exploration, and a reinforcement-learning-like environment are vital.
This isn’t unique to Kimi. Having come from the first generation of “Four Little Dragons,” though they didn’t evolve into the massive companies we now hope for, their atmosphere was remarkably similar. Across generations, these top Chinese AI companies have created a remarkably open environment for China’s brightest, most curious, and most idealistic youth. Founders didn’t say “we must do A, B, C”—they offered a broad invitation: “Maybe this is worth exploring.”
A group of exceptionally talented, highly exploratory individuals in a good environment will inevitably spark innovation. I sense this in both the old-generation Yitu, Megvii, SenseTime, and the new generation—DeepSeek, Kimi, Zhipu. If you place great young minds in a good environment, they will collide and produce something. The core task for these AI companies is to create that environment and constantly adjust the reward system to ensure a robust feedback loop.

Chen Xi: We’ve discussed the first generation of China’s “Four Little Dragons.” Let’s systematize it. The first generation—SenseTime, Megvii, Yitu, CloudWalk—was centered on computer vision. Where were you then?
Kiwi: I joined Yitu in 2016 to work on computer vision. AI 1.0 wasn’t just CV—there were machine learning companies like Fourth Paradigm, autonomous driving, and NLP startups.
But overall, the first generation faced a collective challenge: insufficient model generalization. For example, facial recognition in public security was a core application for the “Four Little Dragons.” When we began selling to police departments across provinces in 2017, we found northeastern regions had more snow, rain, fog—requiring re-labeling and retraining. Another province had vast geography—wider camera coverage, different mounting heights and angles—again necessitating model retraining.
Our major dilemma: AI significantly improved public safety, but poor generalization forced massive production-side services. These services were handled by China’s smartest, most expensive STEM graduates. This turned the business model into a B2B service-type enterprise—like premium outsourcing. But we used highly skilled, expensive talent, making the business model unsexy.
Chen Xi: The market ultimately wasn’t large enough—unlike today’s AGI, lacking imagination.
Kiwi: Exactly. Back then, our models lacked generalization, so every downstream scale required additional R&D investment. Unlike internet companies building standardized systems where marginal cost drops to near zero with scale, we lacked such a model. We were essentially a service company—bigger business volume meant higher costs and increasing management complexity. The underlying business model fundamentally differed from today’s large model companies.
Chen Xi: What did you do at Yitu?
Kiwi: I was a product manager. It’s ironic—joining Yitu in late 2016, the next day, my co-founder Lin Chenxi called the three new hires into a small dark room and announced: “From today, Yitu’s product team is formed.” He assigned each of us a product line and introduced us to our researchers and engineers—then left us alone, pure RL-style.
I remember the first year was extremely painful and stressful—no one told us the correct answers, no guidelines. We made countless mistakes. But two years later, reflecting on it, I’m truly grateful for that AI company. They genuinely trusted young people—both previous and current waves. A core trait of these excellent AI companies: they don’t teach you upfront, but give space to explore. You receive feedback and self-iterate—even if inefficient.
At the time, I envied fresh grads at big tech—knowing processes and standards, avoiding risks through workflows. But we entered a world where no one told us what to do. What did product managers do? Anything researchers and engineers didn’t. We didn’t know the right way—each mistake taught us: “Maybe we should add a checkpoint here,” “Add a review step,” “Align this task.” Every improvement came from learning through failure.
After leaving Yitu, I realized this imprint shaped me deeply. The capability gained through reinforcement has generalization—it’s not learned from someone else. When told to do something, you may not question why. But every process and organizational structure we adopted arose from discovering the need for it through exploration and discussion.
Another point: Yitu had its first official product manager in 2016, though founded in 2012. From 2011 to 2014, SenseTime, Megvii, Yitu and others emerged—one by one—after seeing breakthroughs like ImageNet and AlexNet, believing computer vision could move from academia to impact real life and enter industry.
But honestly, no one had the right answer. It was a research problem—we didn’t know if there was one. So the exploration space was possibly even larger than today—founders didn’t know the answer either. We were just trying to find it—a truly exciting process.
Chen Xi: That’s quite interesting—during the first wave of Chinese AI, people were already boldly hiring newcomers, including recent college graduates.
Kiwi: Even students not yet graduated. I recall many dropouts—some were encouraged to leave school. Others were high school interns. Smart kids came in and explored.

Chen Xi: You mentioned talent flowing from the first generation to today’s “Four Little Dragons” and “Six Little Tigers.” But you also said many were emotionally scarred from past failures and reluctant to re-enter the arduous path of AI innovation or AGI creation. How do these two streams of talent choose?
Kiwi: I think the legacy spans beyond this generation. Looking back at the first “Four Little Dragons,” their predecessors were institutions like Baidu’s AI Research Institute and MSRA—earlier Chinese AI hubs. The core legacy is faith in technology—driven by curiosity, not utilitarianism.
Talent is passed from one generation to the next. I vividly remember during my undergrad years—CS majors or ACM competitors were drawn to MSRA internships. Even Yitu’s co-founder Lin Chenxi briefly interned at MSRA before being mentored by Wang Jian to help launch Alibaba Cloud.
MSRA and Baidu’s AI Research Institutes nurtured many founders of China’s first-gen AI 1.0 era. These founders then nurtured the next wave—founders of AI 2.0. For example, I connected with Zhilin through friends from Megvii and Yitu, whose co-founders included people from Megvii. MiniMax’s founding members include many from SenseTime and Yitu—this generation.
Even today’s model company founders often interned at the “Four Little Dragons.” When I later spoke with researchers and engineers at Zhipu or DeepSeek, I realized we shared a common memory—though we never met, we explored similar paths in that era.
Chen Xi: The community—how close are the relationships? It feels like many are Tsinghua, Fudan, Zhejiang University alumni—often senior classmates, mentors and students—relationships seem tight. When Moonshot was founded, Yang Zhilin chose many of his Tsinghua peers as co-founders.
Kiwi: I think it wasn’t intentional. Fundamentally, idealists coming together to pursue something uncertain require chemical reactions in daily interactions. You just click—feel like you can explore together. It’s influenced by environment—likely groups of brilliant people from SJTU ACM, Tsinghua CS or Yao Class. Even today’s large model wave includes many overseas students—people with technical ideals often share similar environments or competed in ACM.

Chen Xi: You mentioned inheritance. Recently, I saw a striking clip—Professor Tang Jie from Zhipu introducing Yang Zhilin on stage, while Tang was Zhilin’s undergraduate advisor. Now, Zhipu and Moonshot are becoming competitors—when K3 launched, Zhipu’s stock plummeted. Do you think Professor Tang is bothered by this? What’s their relationship like?
Kiwi: First, they’re not competitors. I see them as fellow explorers advancing together. The market is vast. At this stage, even if we ignore products and focus only on model capability, selling model capacity is the future’s intellectual infrastructure. This massive infrastructure cannot be monopolized by one player.
So fundamentally, we face the same challenge: skepticism from non-AI circles or those distant from AI. But in reality, we’re all partners in exploring AGI.
Before ChatGPT, Professors Tang and Zhilin participated in GLM. At Tsinghua, Professor Liu Zhiyuan led the CPM team. Everyone believed large models were worthwhile—yet they demand massive resources. That’s why they founded Beijing Academy of Artificial Intelligence, training China’s first trillion-parameter models—GLM, CPM, and more. Before ChatGPT, they already gave young people ample room to explore with relatively abundant resources. After ChatGPT’s release, everyone rushed to industrial-scale large model training—this laid a fantastic foundation.
Chen Xi: The young pioneers from the first AI 1.0 “Four Little Dragons” are now middle-aged. They’re undoubtedly more experienced—what roles do they play in this round?
Kiwi: They’re all backbone figures. Many don’t appear publicly, but they’re steadfastly pushing things forward. Having lived through cycles, they’re better at filtering out short-term reward noise. Their impact may not be visible to public markets, but they’re driving critical work.
Chen Xi: Roughly speaking, what percentage of that cohort remains in this AI wave? Over 10%?
Kiwi: It depends on scope. Among my friends, over 10% have already joined—some even after 2023. They may not be founders or public faces like Yang Zhilin, but they’re leaders driving key initiatives.

Chen Xi: Let’s talk about DeepSeek—it truly emerged as a dark horse, nobody predicted it.
Kiwi: Indeed, a true dark horse. Honestly, I didn’t notice DeepSeek in early 2023—I didn’t even know they were working on this.
My main information sources were friends from the first AI 1.0 generation—where they went, what they did. DeepSeek only became apparent to me by late 2023, as they hired younger talent. The middle-tier talent from the “Four Little Dragons” mostly joined companies like Moonshot or Zhipu, while DeepSeek recruited younger people. I gradually encountered DeepSeek through deeper engagement with newer researchers and engineers.
When we talked to Liang, his hiring strategy was simple: recruit the smartest people from coding competitions and ACM. Just hire the brightest. So he had no historical baggage—no requirement for experience—directly bringing in exceptional young talents.
Chen Xi: Everyone loves hiring the smartest people. How does this differ from what Yang Zhilin wanted?
Kiwi: The standards are similar. The difference lies in the earliest founders’ backgrounds, affecting how efficiently they reach different talent pools. DeepSeek, originally from Huashu Quantitative, targeted the same pool—ACM and NOI enthusiasts—because quant firms offer high cost-efficiency. This gave DeepSeek natural access to young, top-tier NOIs and ACM competitors.
Meanwhile, MiniMax and Kimi’s co-founders came from the “Four Little Dragons,” so they accessed talent from that supply chain. Similarly, Zhipu, born from Tsinghua labs, has strong reach within Tsinghua’s recent graduates. So the difference isn’t necessarily in talent taste—but in the earliest co-founders’ identities and environments, affecting their talent acquisition efficiency and distribution.

Chen Xi: After leaving Yitu, you went to Kai-Fu Lee’s firm.
Kiwi: Yes, I joined Sinovation’s AI Engineering Institute. Many don’t know—Sinovation is a VC, but our AI Engineering Institute once peaked at 200+ people.
I joined in 2018, but the institute was established earlier—around 2016–2017, around the time I joined Yitu. I later asked the director, Wang Yonggang, why Sinovation created the AI Engineering Institute.
Simply put, Kai-Fu Lee and Wang Hua were among the earliest Chinese VCs to heavily invest in AI. In the AI 1.0 era, Sinovation backed nearly all major sub-sectors—“Four Little Dragons,” Fourth Paradigm, Horizon Robotics (autonomous driving)—making Sinovation one of the earliest, highest-commitment funds in China.
Yet, during this journey, we clearly identified flaws in the AI 1.0 paradigm. The problems I saw at the “Four Little Dragons” mirrored those we observed during investment: scientists-turned-founders lacked commercial training and experience. We served traditional enterprises—ultimately doing consulting, similar to software or SaaS.
Crucially, SaaS faced huge resistance in China—modern corporate development is only ~20–30 years old, with no fully professional management systems. We believed in management best practices, but entrepreneurs were still exploring—no standardized blueprint existed.
Thus, in AI 1.0, regardless of niche or downstream clients, commercial scaling resembled traditional software services—making the journey painful. We never found a viable path like the internet’s model.
So the AI Engineering Institute’s premise was: if all are B2B, all AI, can we do it more efficiently—helping with product development, commercialization, and scaling projects? When I joined, there was no clear job title. Each stage’s role was ambiguous.
My role was vaguely defined as “Frontier Domain Exploration Lead.” Our core task: regularly track academic breakthroughs in each AI niche. For example, we experienced GANs, federated learning, and GPT. After breakthroughs, we assessed if real-world production or lifestyle scenarios could apply these algorithms—doing 0-to-1 product demo exploration.
If proven viable, we’d assess market size and consider incubating new AI companies. Sinovation’s AI Engineering Institute thus incubated Sinovation Intelligent, which later went public. My final venture post-Sinovation was LanZhou Tech—incubated after GPT-2 and GPT-3’s releases. Throughout, we were exploring how AI could penetrate real human life and production.
Chen Xi: When did you realize the scaling law phenomenon began in Silicon Valley?
Kiwi: I’m embarrassed to admit—only after ChatGPT. But in 2019, we attempted something: “Chinese GPT-2.” Inside Sinovation’s AI Engineering Institute. We studied GPT-2’s paper—needed data, GPUs, NLP experts. Ironically, Kai-Fu is a master connector—many friends from his network paused at Sinovation during career transitions.
When Tencent AI Lab’s Zhang Tong and his NLP researchers left Tencent, they briefly joined us—giving us qualified NLP researchers. Around summer, we got thrilling news: one of our Middle Eastern LPs discovered hundreds of unused V100s in a warehouse—dormant for months.
Since Sinovation’s strong public image is AI, they probably saw our portfolio and assumed we were GPU-related. We don’t know why they bought, but they approached Kai-Fu: “We found hundreds of idle V100s—do you want them?” We said yes—so we had GPUs.
With GPUs and people, data wasn’t scarce in China. We began training. After two months, I tested it—results were terrible. Asked to reply to a simple email, it output three commas consecutively. Honestly, we felt disappointed.
Our post-mortem: Chinese data quality wasn’t as high as English corpora. But by 2020, after another attempt, I realized I hadn’t correctly diagnosed the issue. Our core team consisted of researchers, but the critical factor was data—data needed thorough cleaning and processing. We lacked manpower for proper data engineering, leading to poor data quality and subpar model performance. Also, we hadn’t grasped scaling law fundamentals—never read the seminal paper.
So in 2020, when we tried incubating LanZhou Tech again, we consciously paired industrial veterans with extensive dirty work and engineering experience to complement Professor Zhou Ming. Zhou is a respected deputy director at MSRA, with 30 years in NLP research.
That time, we brought more engineers and emphasized data. But the core problem remained: my understanding of scaling law was inadequate. We still applied the old “Four Little Dragons” model—seeking vertical B2B applications. In short, I carried a strong failure path dependency from AI 1.0: selling model capability isn’t a good business.

Chen Xi: But now, it’s a necessary business—Model as a Service.
Kiwi: I’m uncertain—this may be a non-consensus view. As someone full of bias, I hold a bold opinion: selling model capability isn’t a good business. Even if Claude Code proves rapid revenue growth, I believe Claude Code is a great product. It’s not just model capability—it layers a powerful interface atop the model, providing programmers an immediate feedback environment in coding production. It includes numerous harnesses and tools. Every task and instruction to AI gets instant feedback: “Should I use this solution?” “Do I need debugging?” This creates a high-quality feedback loop—excellent signal.
Thus, I believe model capability—whether API or SDK—is destined to commoditize. That’s a conviction.
I’m not saying commoditized businesses aren’t good. Today’s cloud computing—several cloud vendors—represents a solid business. But overall, it doesn’t generate high excess profits. It’s the next phase’s intellectual infrastructure—a land-grabbing game. You must first build this infrastructure, secure the land, then layer better environments—rich in feedback, context, and reward—to explore further tasks.
Next, more bold opinions—my personal biases. I see a strong market narrative: “Model is all.” I believe models will gradually absorb consensus, in-distribution tasks—including the harness layer. But I also believe human intelligence cannot be exhausted—human society’s exploration of intelligence is endless.
So, returning—what will truly emerge as the Super App in this AGI wave? I don’t know what it looks like, but it will be an environment where continuous interaction between humans and AI generates new intelligence. That’s AGI’s future—not merely selling model capability. If AGI’s entire product form stops at selling model capability, it won’t be the exciting prospect we envisioned in late 2022 and early 2023. Because I believe human intelligence cannot be exhausted.
Chen Xi: A few years ago, there was a VC assumption: “winner takes all.” Make the strongest model, stay first forever, earn the most money, the most profit. But now it seems untrue—no one guarantees perpetual leadership, and the gap between first and second/third place isn’t that wide.
Kiwi: This depends on the stage. If we look purely at selling APIs or model capability, this wave differs little from the first “Four Little Dragons”: during rapid model development, before stability, SOTA models control pricing, yielding super profits. Followers can’t achieve super profits—must sell near cost or even at a loss.
But from our experience with the first “Four Little Dragons,” maintaining technical leadership is impossible. Technical edge offers only a window—“you have what others don’t.” Crucially, during this window, can you build a business or product barrier?
For example, my expectation for this AGI product: when you have what others lack, can you build an environment that continuously generates new intelligent data—where human-product interaction fuels ongoing model iteration? If successful, more intelligence accumulates in your ecosystem. The flywheel spins upward—others can’t catch up. If not, the window closes—you become commoditized, all model capabilities equalized.

Chen Xi: This question might be blunt—I’m just curious. Sinovation touched many cutting-edge insights early, had GPUs—why didn’t Kai-Fu’s ZeroOne succeed?
Kiwi: First, ZeroOne is still developing superbly. Recently, they’ve been doing many FDE (Forward Deployment Engineer) projects.

Chen Xi: I know—recently they aim to become China’s Palantir.
Kiwi: Yes. I think FDE isn’t easy—it’s not purely tech-driven. Core to FDE: convincing non-AI professionals to accept AI integration. Frankly, AI is highly centralized—people’s instinctive reaction is fear. Knowledge workers fear AI will distill them—teach AI their job, then replace them.
FDE implementation requires combining business acumen, organizational understanding, human interaction, technical insight, and product sense. I believe Kai-Fu is perfectly suited for FDE—he’s been evangelizing AI for years. So ZeroOne is doing valuable work—still deeply committed.
Chen Xi: Earlier, when competing on model capability, what did ZeroOne lack?
Kiwi: First, Kai-Fu, Wang Hua, and ZeroOne have always deeply invested in AI and consistently pursued it. But new ventures need people free of burdens. Honestly, talent performs differently in different environments—even the same person varies. For example, my demeanor differs in a podcast interview vs. a private dinner.
So placing the same talent in different environments yields different results. What’s the ideal environment for large models? One where “pursuing AGI” and “proving you can pursue AGI” aren’t seen as vastly different. This requires relatively frontline management. If management is distant, in my view, rewards may drift.

Chen Xi: Chinese tech giants—ByteDance, Alibaba, Tencent—played crucial roles in this AI wave. How do their roles differ from Moonshot, Zhipu—newborn, AI-native model companies?
Kiwi: Quite different. For example, ByteDance surprised me. This year, Seedance 2.0, AI short dramas, combined with TikTok’s content incubation platform, created a powerful closed-loop model: model production, consumption, and distribution in AGI iteration. This loop may even surpass Claude Code’s in advancement.
So ByteDance, a company that mastered ecosystem strategies in mobile internet, combines its strategic awareness with massive, cost-insensitive AGI investment—yielding many innovative plays.
Alibaba: Qwen is an excellent team. Alibaba has invested in nearly all major Chinese model companies—providing massive GPU resources. So Alibaba is effectively a land-grab supporter—building the foundational infrastructure for China’s entire large model ecosystem—critically important.
Tencent: Though recent model launches haven’t entered the top tier, I reiterate: for a great large model product, the environment is crucial. WeChat possesses one of the best environments I’ve seen in human society. So Tencent remains an unstoppable force in future large model products.
Chen Xi: I see Meituan recently led Moonshot’s latest round—Xing Ge still seems satisfied with Kimi. Last round they didn’t invest, but this round they lead again.
Kiwi: I don’t know the details—I’m now fully focused on building our “wrapper product.”
Chen Xi: When did you leave Meituan? Are you starting your own venture?
Kiwi: Yes, I left Meituan at the end of last year.
I personally believe the next generation of wrappers—our “wrapper products”—have a crucial task: democratizing AI for ordinary people. I still feel deeply that since late 2022, using ChatGPT to today, I’ve been an “inferior citizen” in the AI world.
Those who excel at crafting prompts, building agent loops, mastering AI—those who wield AI well—are “first-class citizens.” But in an open input box, I feel I’m not smart enough. Same question—someone with engineering experience can build a more solid development task; I might get “shitcode” from AI—unmaintainable.
So I’ve consistently felt like an “inferior citizen” of AI products—still do. But if a product makes a segment of users feel inferior, it’s likely not a good product—or not consumer-friendly. I believe the next phase must focus on democratizing intelligence—making it accessible and equitable for all. That’s why I see immense potential in the next wave of “wrapper products.”
Chen Xi: Is it convenient to disclose what you're working on right now?
Kiwi: We can share a bit—right now we're building a human-centric interaction tool, focused on how to better serve interpersonal communication. This is something we envision every individual using, because everyone experiences life differently and holds unique perspectives on people. We believe human-to-human communication isn't merely about encoding a signal and sending it to you for decoding—it's about co-creating an interactive space where intent is reconstructed together. So in the near future, we may release a lightweight wrapper tool built around this concept for public use.
Chen Xi: Sounds great—I'm really looking forward to your product launch. Thank you, Kiwi.
Kiwi: Thank you.
Source: Silicon Valley 101
Disclaimer: Contains third-party opinions, does not constitute financial advice
AI drives SK Chairman's divorce settlement bill to KRW 94.4 billion
4 days ago
AI is no longer competing on benchmark scores, but on profitability
4 days ago
ByteDance's Douyin Beans priced at 68 RMB—worth it for 382 million monthly active users?
4 days ago
Fields Medalist Concerned About AI Extinction Heads to OpenAI
4 days ago
30 Million KRW Threshold, AI Chip Leverage Cooling Down
5 days ago
Meta gives away models for free—who dares to price AI now?
5 days ago
$13.7 Billion Prediction Market: Are Retail Investors Making Way for AI?
5 days ago






