General text & conversation
Live · all tiers · default
Every Bee model handles text generation, conversation, writing, summarisation and streaming responses. This is the baseline, not Bee's product category.
Bee Models
Bee includes general text and conversation, but it is built for specialist work: reasoning, research, analysis, coding, tools, documents and governed workflows. Bee Cell through Bee Swarm are controlled model lines spanning 128K to 1M context. Every alias resolves to a governed production configuration, and an adapter can serve only after its held-out evaluation and release checks pass. Enclave is a private deployment contract. Bee Ignite is the active HEOSSI foundation research programme, but it is not a customer-selectable model and has no released checkpoint today.
Bee release system
Cell through Swarm are stable Bee product and API identities. Each release binds its serving configuration, specialist capabilities, safety policy, evaluations and immutable release evidence through the Bee Specialist Intelligence System.
Bee's hosted inference is operated by HEOSSI. Its serving configuration, governance, safety enforcement, routing, retrieval, tools, evaluations, and release controls are designed, implemented, and operated by HEOSSI through the Bee Specialist Intelligence System (BSIS).
Bee Ignite foundation programme
HEOSSI is developing the research, provenance and training control plane required for a future independently pretrained specialist foundation model through the active Bee Ignite foundation research programme. Ignite is not a transfer or adapter lane and has no released checkpoint today.
Explore BSISModel specifications and product evidence
Generated from the pinned release catalog and the live-verified capability source of truth. Last capability verification: 2026-07-14. Higher plans may route a task to an included lower lane; this table describes each named primary lane rather than the full set of services available in the product.
| Bee tier | Architecture & serving | Context | Text & reasoning | Native image | Native video | Tool-call API | Strict JSON | 1M |
|---|---|---|---|---|---|---|---|---|
| Bee Cell Free, fast multimodal entry tier | 4B dense · unified vision-language Bee-hosted vLLM lane | 128K native | Live | Live | — | Live | — | — |
| Bee Brood Cost-efficient multimodal reasoning | 9B dense VL · thinking-enabled Bee-hosted vLLM lane | 256K native | Live | Live | Live | Live | — | — |
| Bee Comb Structured multimodal workhorse | 27B dense VL · FP8 Bee-hosted vLLM lane + 1M twin | 256K native · 1M routed | Live | Live | Live | Live | Live | Live |
| Bee Buzz Agent and builder workhorse | 35B-A3B sparse MoE VL · FP8 Bee-hosted vLLM lane + 1M twin | 256K native · 1M routed | Live | Live | Live | Live | Live | Live |
| Bee Hive High-capability specialist intelligence | 122B-A10B sparse MoE VL · NVFP4 Bee-hosted vLLM lane | 256K native | Live | Live | Live | Live | Live | — |
| Bee Swarm Premium distributed intelligence | Versioned multi-lane routing policy Governed routing fabric | 1M native | Live | Routed | — | — | — | Live |
Core intelligence and workflows
General text and chat provide the baseline. Bee is differentiated by specialist reasoning, current-source research, coding and tools, document intelligence, durable workspaces, automation and governed long-context work.
Live · all tiers · default
Every Bee model handles text generation, conversation, writing, summarisation and streaming responses. This is the baseline, not Bee's product category.
Live · all tiers · tier-scaled
Bee is designed for technical, business, research and governed specialist workflows. Reasoning depth, context and assurance increase through the tiers.
Live · all tiers · selectable
Web Search retrieves current sources with citations. Research mode plans multiple searches, reads across results and synthesizes a sourced answer.
Live · tier-aware
Coding assistance is available across the lineup. Structured tool calls are exposed through the lanes proven in the model matrix above.
Live · all tiers
Projects organise ongoing work, memory preserves useful context, and scheduled tasks run prompts and deliver notifications.
Live · all tiers
Tenant-scoped retrieval, full-document injection when it fits, and citations ground specialist work in the user's own source material.
Live · automatic + manual
Every tier compacts at 75% of its context budget. Recent multimodal turns stay verbatim; older media is represented explicitly in the summary.
Supporting modalities and specialised compute
These services let specialist workflows accept and produce more than text. They extend Bee’s intelligence layer; they do not turn Bee into a media stack.
Live · all tiers
Speech-to-text and text-to-speech are provider-routed Bee services available from Cell through Swarm.
Live · multimodal composer
The workspace accepts images, audio, video, PDFs, office files, code and text. Native video analysis remains limited to the lanes proven above.
Live · plan-capped
Text-to-image generation is a prepaid, governed service separate from image understanding: a 3/day promotional safety ceiling on Cell and progressively higher funded-plan ceilings.
Built · not active
The generation workflow is not sold as live until completed-job URL delivery passes production verification.
Opt-in · entitlement-gated
Simulator and OpenQuantum hardware jobs use a dedicated endpoint. Ordinary chat never silently routes through a QPU.
Stable model names across the workspace and OpenAI-compatible API. Availability follows your plan and current serving capacity.
bee-cell · 128K context · base training-data cutoff not publicly disclosed
Current-source Web Search and Research available
Input
$0.30 / 1M
Output
$1.20 / 1M
Best for
Capabilities
bee-brood · 256K context · base training-data cutoff not publicly disclosed
Current-source Web Search and Research available
Input
$0.60 / 1M
Output
$2.40 / 1M
Best for
Capabilities
bee-comb · 256K context · base training-data cutoff not publicly disclosed
Current-source Web Search and Research available
Input
$1.25 / 1M
Output
$5.00 / 1M
Best for
Capabilities
bee-buzz · 256K context · base training-data cutoff not publicly disclosed
Current-source Web Search and Research available
Input
$2.50 / 1M
Output
$10.00 / 1M
Best for
Capabilities
bee-hive · 256K context · base training-data cutoff not publicly disclosed
Current-source Web Search and Research available
Input
$5.00 / 1M
Output
$20.00 / 1M
Best for
Capabilities
bee-swarm · 1M context · base training-data cutoff not publicly disclosed
Current-source Web Search and Research available
Input
$10.00 / 1M
Output
$40.00 / 1M
Best for
Capabilities
A private training candidate is not a public model release. Bee promotes an adapter only after held-out evaluation, independent rubric comparison, and a reproducible snapshot are recorded. Validation criteria and release evidence are published at /trust.
Under the hood
The six public tiers are exposed as pinned, controlled Bee releases. Each tier uses a stable alias, preserves an immutable serving revision, and adds optional per-domain adapters only through an eval-gated release pipeline. That keeps multimodal, reasoning, tool, and long-context capability tied to reproducible evidence instead of a floating model label.
Capabilities
Difficulty-scored routing between local execution and frontier-teacher escalation. Scoring blends keyword complexity, query length, conversation depth, code/math detection, and a per-domain multiplier.
source · Adaptive router
The evolution loop generates candidate neural modules — attention variants, SSM discretisations, compression codecs, memory protocols — runs them through a sandboxed eval, and only accepts winners. Architecture only; not currently running against the live Bee Cell deployment.
source · Evolution engine
When a request needs code, the self-coding module is designed to write it, execute in a sandbox, read the error, and iterate. It is enabled only for tiers and workflows that have cleared their release checks.
source · Self-coding & self-healing
Document upload → chunk → embed → FAISS index → cite. End-to-end retrieval is built into every tier; pass document IDs in the chat request and Bee handles the rest.
source · RAG pipeline
Bounded real-QPU execution through OpenQuantum, with a local statevector simulator for development. Chat requests are never routed through quantum; Quantum Hardware Pack holders run explicit, metered jobs through the dedicated endpoint.
source · Quantum integration
LoRA adapters specialise the base model for low-cost fine-tuning per domain. Adapters are released through the governed-release pipeline once each clears the eval harness — every released adapter has a published validation record at /trust.
source · Per-adapter validation records published at /trust
API compatibility
Bee speaks the same wire protocol as OpenAI's /chat/completions. Switch the base URL and the API key — your existing client code keeps working. Streaming is universal; tools, JSON mode, structured outputs, vision, and video are exposed on the tiers whose live model cards list them.
bee-enclave · contracted · knowledge boundary follows the selected governed lane
Customer-scoped Hive- and Swarm-class deployment designs can cover private, regulated, or sovereign requirements. Broader self-hosted and air-gapped operation remains roadmap.
Research work is active. Model serving is not. Bee Ignite becomes a model only after a HEOSSI-pretrained checkpoint clears the full BSIS release process.
bee-ignite · HEOSSI foundation programme · pretraining not started
HEOSSI-internal foundation-model programme inside BSIS. Bee Ignite requires a HEOSSI tokenizer, rights-qualified corpus, random-initialized architecture, reproducible distributed pretraining, specialist evaluation, shadow, canary and rollback evidence. No checkpoint is currently released or commercially available. Earlier transfer, adapter, and experimental serving lanes are retired; they do not establish independent foundation-model provenance.
Now
Programme controls and provenance
Next
Scaling, architecture and tokenizer
Then
Corpus and distributed pretraining
Release
Evaluation, shadow, canary and rollback
Programme workstreams