Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

Fireworks Releases Ember-1 Efficient Model

Fireworks Releases Ember-1 Efficient Model

Fireworks AI·Thursday, September 24, 2026
  • •Fireworks Research released Ember-1, built on Kimi K3, claiming comparable quality with 40% fewer tokens.
  • •Ember-1 set a cost-per-task frontier on Bedside Bench, covering 500 clinical cases across 10 specialties.
  • •Two customer coding tests found approximately 35% fewer tokens per task at comparable quality; one customer deployed it.
  • •Fireworks Research released Ember-1, built on Kimi K3, claiming comparable quality with 40% fewer tokens.
  • •Ember-1 set a cost-per-task frontier on Bedside Bench, covering 500 clinical cases across 10 specialties.
  • •Two customer coding tests found approximately 35% fewer tokens per task at comparable quality; one customer deployed it.
  • •Fireworks Research released Ember-1, built on Kimi K3, claiming comparable quality with 40% fewer tokens.
  • •Ember-1 set a cost-per-task frontier on Bedside Bench, covering 500 clinical cases across 10 specialties.
  • •Two customer coding tests found approximately 35% fewer tokens per task at comparable quality; one customer deployed it.
  • •Fireworks Research released Ember-1, built on Kimi K3, claiming comparable quality with 40% fewer tokens.
  • •Ember-1 set a cost-per-task frontier on Bedside Bench, covering 500 clinical cases across 10 specialties.
  • •Two customer coding tests found approximately 35% fewer tokens per task at comparable quality; one customer deployed it.

Fireworks Research released Ember-1 on September 23, 2026, as a specialized model built on Kimi K3. Fireworks says it matches Kimi K3’s quality with 40% fewer tokens by cutting unnecessary reasoning. The company tested it on public benchmarks, customer production traffic and internal coding workloads, and said quality held across those settings. Ember-1 is available as a Research Preview on Fireworks Serverless alongside the base Kimi K3 model.

Fireworks says users wanted Kimi K3’s coding capabilities at lower cost because long reasoning traces made automated coding expensive at scale. Lowering K3’s reasoning-effort setting reduced quality, so Fireworks trained a model to reason more efficiently. Its team ran more than 50 training experiments and over 200 evaluations using Fireworks Serverless Training, which let it run experiments without provisioning or managing GPUs. Training covered mathematics, coding, instruction following, conversation, search, tool use and software engineering, including standalone tasks and extended interactions.

Fireworks says reasoning can account for more than 90% of tokens generated by models such as Kimi K3. In multi-turn agent workloads, earlier reasoning is passed back into each later call, causing context to grow roughly quadratically with the number of turns. The company says Kimi K3’s reasoning is often longer than tasks require, though self-reflection—such as revisiting assumptions or using feedback to recover from mistakes—can be useful. Across seven benchmarks and production traffic from two customers, Fireworks reports that Kimi K3’s reasoning was shortened by 35–50% without sacrificing accuracy.

On Doximity’s Bedside Bench, a physician-validated benchmark of 500 clinical cases across 10 specialties, Ember-1 set a new cost-per-task Pareto frontier among open and closed models, including GPT-5.6 Sol, GPT-6 Astra and Claude Opus 5. Fireworks also compared Ember-1 with Kimi K3 at low, high and maximum reasoning effort, using public Kimi K3 API rates: $3 per million uncached input tokens, $0.30 per million cached input tokens and $15 per million output tokens. On benchmarks with more than 50 samples, Fireworks said Ember-1 was on or near the cost-quality frontier, matching K3-max quality at lower cost and outperforming K3-low.

In the reported benchmark results, Ember-1 scored 82.0% on Terminal Bench 2.1 (89 samples), versus 80.9% for K3-max, with 51.9% lower cost, a $23.1 saving. On SWE-bench Verified (500 samples), it scored 92.2%, versus 93.2%, at 15.5% lower cost, a $68.1 saving. On SWE-Interact (75 samples), scores were 20.0% versus 21.3%, with 32.5% lower cost and a $60.8 saving. On DeepSWE 1.1 (113 samples), Ember-1 scored 75.2% against 66.4%, with 23.7% lower cost and a $126.9 saving. On τ-2 Bench Airline (50 samples), it scored 66% against 64%, with 5.9% lower cost and a $0.3 saving.

In live A/B tests on two customers’ production coding workloads, Fireworks reported approximately 35% fewer tokens per task at comparable quality. One customer has moved Ember-1 into production and plans to scale it to replace the base model. Fireworks also said its developers used Ember-1 for everyday coding without noticing the internal switch: its reported score was 0.753, compared with 0.750 for Kimi K3, while reasoning made up 71.3% of tokens and total token use fell 34.5%.

Fireworks is offering Ember-1 for two weeks through Serverless as a research release, with permanence depending on community demand. The company is also launching training support so enterprises can customize Ember-1 with their own data. Fireworks says it plans to develop additional specialized Ember models for workloads where reasoning tokens make up most of the cost.

Fireworks Research released Ember-1 on September 23, 2026, as a specialized model built on Kimi K3. Fireworks says it matches Kimi K3’s quality with 40% fewer tokens by cutting unnecessary reasoning. The company tested it on public benchmarks, customer production traffic and internal coding workloads, and said quality held across those settings. Ember-1 is available as a Research Preview on Fireworks Serverless alongside the base Kimi K3 model.

Fireworks says users wanted Kimi K3’s coding capabilities at lower cost because long reasoning traces made automated coding expensive at scale. Lowering K3’s reasoning-effort setting reduced quality, so Fireworks trained a model to reason more efficiently. Its team ran more than 50 training experiments and over 200 evaluations using Fireworks Serverless Training, which let it run experiments without provisioning or managing GPUs. Training covered mathematics, coding, instruction following, conversation, search, tool use and software engineering, including standalone tasks and extended interactions.

Fireworks says reasoning can account for more than 90% of tokens generated by models such as Kimi K3. In multi-turn agent workloads, earlier reasoning is passed back into each later call, causing context to grow roughly quadratically with the number of turns. The company says Kimi K3’s reasoning is often longer than tasks require, though self-reflection—such as revisiting assumptions or using feedback to recover from mistakes—can be useful. Across seven benchmarks and production traffic from two customers, Fireworks reports that Kimi K3’s reasoning was shortened by 35–50% without sacrificing accuracy.

On Doximity’s Bedside Bench, a physician-validated benchmark of 500 clinical cases across 10 specialties, Ember-1 set a new cost-per-task Pareto frontier among open and closed models, including GPT-5.6 Sol, GPT-6 Astra and Claude Opus 5. Fireworks also compared Ember-1 with Kimi K3 at low, high and maximum reasoning effort, using public Kimi K3 API rates: $3 per million uncached input tokens, $0.30 per million cached input tokens and $15 per million output tokens. On benchmarks with more than 50 samples, Fireworks said Ember-1 was on or near the cost-quality frontier, matching K3-max quality at lower cost and outperforming K3-low.

In the reported benchmark results, Ember-1 scored 82.0% on Terminal Bench 2.1 (89 samples), versus 80.9% for K3-max, with 51.9% lower cost, a $23.1 saving. On SWE-bench Verified (500 samples), it scored 92.2%, versus 93.2%, at 15.5% lower cost, a $68.1 saving. On SWE-Interact (75 samples), scores were 20.0% versus 21.3%, with 32.5% lower cost and a $60.8 saving. On DeepSWE 1.1 (113 samples), Ember-1 scored 75.2% against 66.4%, with 23.7% lower cost and a $126.9 saving. On τ-2 Bench Airline (50 samples), it scored 66% against 64%, with 5.9% lower cost and a $0.3 saving.

In live A/B tests on two customers’ production coding workloads, Fireworks reported approximately 35% fewer tokens per task at comparable quality. One customer has moved Ember-1 into production and plans to scale it to replace the base model. Fireworks also said its developers used Ember-1 for everyday coding without noticing the internal switch: its reported score was 0.753, compared with 0.750 for Kimi K3, while reasoning made up 71.3% of tokens and total token use fell 34.5%.

Fireworks is offering Ember-1 for two weeks through Serverless as a research release, with permanence depending on community demand. The company is also launching training support so enterprises can customize Ember-1 with their own data. Fireworks says it plans to develop additional specialized Ember models for workloads where reasoning tokens make up most of the cost.

Read original (English)·Sep 23, 2026
#ember 1#fireworks research#kimi k3#token efficiency#reasoning models#bedside bench#swe bench verified#serverless training