Korean chip startups are betting that inference costs decide the AI market
A small group of Korean fabless designers is aiming at the recurring cost of serving AI models rather than the one-off cost of training them, on the theory that the second market is bigger and less defended.

Nobody sensible starts a company to displace Nvidia at training large models. The capital requirements are enormous, the software ecosystem is a decade deep, and the customers are the handful of firms best equipped to negotiate. The Korean AI-chip startups — Rebellions and FuriosaAI are the visible examples of the category, alongside the neural-processor effort that SK Telecom spun out as Sapeon — have therefore aimed at a different line on the same budget: inference, the cost of actually answering queries once a model exists.
The argument for that target is an accounting one. Training a model is a capital event, undertaken occasionally and amortised. Inference is an operating cost that recurs with every request, scales with usage, and shows up on an electricity bill. As deployment broadens, total spending shifts from the first category toward the second. Inference workloads are also shaped differently: they are latency-sensitive, run in smaller batches, and are frequently constrained by memory bandwidth and power rather than raw arithmetic throughput. A part that delivers fewer operations per second but substantially more useful output per watt can win procurement even against a technically superior competitor, because the buyer is paying for the watt.
Korea has one structural reason to be plausible here that other places lack, and it is high-bandwidth memory. The stacked DRAM that determines how fast an accelerator can be fed is manufactured by two Korean companies, and designing a chip in the same country as its memory supplier — with access to advanced packaging, co-design conversations and a deep pool of engineers trained by the memory industry — is a real if unglamorous advantage. Korean accelerator startups have accordingly emphasised memory-side efficiency in their designs rather than peak compute figures.
The domestic market has already begun consolidating, which tells its own story. Rebellions and Sapeon agreed during 2024 to combine, a transaction completed around the turn of 2025 with SK’s backing, on the straightforward logic that Korea could not sustain two subscale neural-processor programmes competing for the same engineers and the same customers. FuriosaAI, founded in 2017, introduced its second-generation inference part in 2024 with a similar emphasis on performance per watt for serving transformer models.
The obstacle is not silicon. It is the compiler. An accelerator is only as fast as the software stack that can map a model onto it, and Nvidia’s durable advantage is that essentially every framework, kernel library and research artefact assumes its hardware. A new architecture must support whatever model shape becomes fashionable next quarter — a different attention variant, a new quantisation scheme, a mixture-of-experts routing pattern — or its efficiency is theoretical. Compiler and runtime teams are more expensive than design teams, take longer to build, and produce nothing demonstrable for years. Most accelerator startups that have failed internationally failed here.
The Korean state has positioned itself as the anchor customer. The K-Cloud initiative, launched by the science ministry to place domestically designed neural processors in public-sector data centres, and the national AI computing centre plans set out in 2024, both direct guaranteed early demand toward local hardware. For a chip startup, committed volume before commercial traction is close to the only thing that cannot be raised from investors. The risk is the mirror image: a company that optimises for a procurement specification learns to satisfy an evaluator rather than a market, and public buyers are not known for demanding the fastest-moving software support.
The window these firms are aiming at is genuine and it is narrow. An inference ASIC earns its advantage by baking assumptions about model architecture into hardware. That is a good trade when architectures have stabilised, and an expensive one while they are still moving.