Qwen: Qwen3.8 27BNewTEETrusted Execution Environment
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be enabled or disabled. Served on Phala in a TDX-attested enclave.
- Aug 24, 2026
- 262K context
- $0.40/M input
- $3.00/M output
- $0.15/M cache read
Meta: Muse Glimmer 30BNewTEETrusted Execution Environment
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon agentic and coding workflows, with multi-step reasoning, reliable tool use, failure recovery, image understanding, and multilingual support across more than 100 languages.
- Aug 13, 2026
- 131K context
- $0.30/M input
- $1.10/M output
- $0.04/M cache read
DeepSeek: DeepSeek V4 Flash 0731NewTEETrusted Execution Environment
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.
- Aug 4, 2026
- 1M context
- $0.20/M input
- $0.40/M output
- $0.07/M cache read
MoonshotAI: Kimi K3NewTEETrusted Execution Environment
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows.
- Jul 29, 2026
- 1M context
- $3.00/M input
- $15.00/M output
- $0.30/M cache read
Z.ai: GLM 5.2TEETrusted Execution Environment
GLM-5.2 is Z.ai's flagship model for the era of long-horizon tasks. With a truly usable 1M-token context window, it can handle project-level engineering context and execute long-running tasks more reliably. Served as a text-only TEE deployment via Phala.
- Jun 16, 2026
- 1M context
- $1.26/M input
- $3.00/M output
- $0.22/M cache read
Qwen: Qwen3.6 27BTEETrusted Execution Environment
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities accepting text and image inputs, a configurable thinking/reasoning mode, and a native 262K context window. Served as a TEE deployment via Chutes.
- Jun 4, 2026
- 262K context
- $0.32/M input
- $2.70/M output
- $0.15/M cache read
DeepSeek: DeepSeek V4 FlashTEETrusted Execution Environment
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
- Jun 2, 2026
- 1M context
- $0.20/M input
- $0.40/M output
- $0.07/M cache read
Qwen: Qwen3.5-122B-A10BTEETrusted Execution Environment
Qwen3.5-122B-A10B is a large Mixture-of-Experts model from Alibaba Cloud with 122B total parameters and 10B active parameters per token. Strong on reasoning, coding, and tool calling with 262K context. Served as a text-only TEE deployment via NEAR AI.
- May 26, 2026
- 262K context
- $0.46/M input
- $3.68/M output
Qwen: Qwen3 32BTEETrusted Execution Environment
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a thinking mode for complex tasks and a standard mode for general dialogue. Served as a TEE deployment via Chutes.
- May 26, 2026
- 41K context
- $0.12/M input
- $0.50/M output
- $0.052/M cache read
Google: Gemma 4 31BTEETrusted Execution Environment
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense model. Features a 256K token context window, configurable thinking/reasoning mode, native function calling, and strong multilingual performance. Served as a text-only TEE deployment via NEAR AI.
- May 26, 2026
- 262K context
- $0.15/M input
- $0.46/M output
- $0.075/M cache read
Qwen: Qwen3.6 35B A3BTEETrusted Execution Environment
Qwen3.6-35B-A3B is an open-weight model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated attention. Served as a text-only TEE deployment via NEAR AI.
- May 26, 2026
- 262K context
- $0.20/M input
- $1.27/M output
Phala: Gemma-4 26B-A4B Uncensored (Heretic)TEETrusted Execution Environment
Uncensored "Heretic" variant of google/gemma-4-26B-A4B-it created using Heretic v1.2.0 with the Arbitrary-Rank Ablation (ARA) method and row-norm preservation. Refusals drop from 100/100 to 11/100 with KL divergence 0.0499 vs the base model. The base Gemma 4 26B A4B is a Mixture-of-Experts model with 25.2B total / 3.8B active parameters (8 active / 128 total experts), 30-layer transformer with hybrid local sliding (1024) + global attention, supporting a 256K context window. Natively multimodal (text + images, variable aspect ratios). Strong on coding, reasoning, function calling, with native system prompt support across 35+ languages. Served on Phala in TDX-attested H200 enclave with end-to-end ECDSA response signing; vLLM-compatible FP8-Static quantization by cloud19 (router excluded from quantization).
- May 23, 2026
- 66K context
- $0.15/M input
- $0.70/M output
Phala: Qwen3.6 35B-A3B Uncensored (Aggressive)DeprecatingTEETrusted Execution Environment
Uncensored "Aggressive" variant of Qwen3.6-35B-A3B from Alibaba's Qwen team. The fine-tune by HauhauCS removes refusal behaviors (0/465 refusals) without modifying datasets or core capabilities. The base architecture is a 35B-parameter Mixture-of-Experts model with 256 experts routing 8 per token (~3B active params), 40 layers, and a hybrid linear+full-softmax attention mechanism (3:1 ratio). Supports a native 262K context and is natively multimodal across text, images, and video. Served on Phala in TDX-attested H200 enclave with end-to-end ECDSA response signing; FP8 quantization by lamianlbe.
- May 23, 2026
- 131K context
- $0.30/M input
- $1.50/M output
MoonshotAI: Kimi K2.6TEETrusted Execution Environment
Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and demonstrates strong performance in agentic workflows.
- Apr 21, 2026
- 262K context
- $1.09/M input
- $4.60/M output
- $0.37/M cache read
Z.ai: GLM 5.1TEETrusted Execution Environment
GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...
- Apr 20, 2026
- 203K context
- $1.21/M input
- $4.20/M output
- $0.60/M cache read
Qwen: Qwen3.5-27BTEETrusted Execution Environment
The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B.
- Mar 13, 2026
- 262K context
- $0.30/M input
- $2.40/M output
- $0.15/M cache read
Qwen: Qwen3.5 397B A17BTEETrusted Execution Environment
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers state-of-the-art performance comparable to leading-edge models across a wide range of tasks, including language understanding, logical reasoning, code generation, agent-based tasks, image understanding, video understanding, and graphical user interface (GUI) interactions. With its robust code-generation and agent capabilities, the model exhibits strong generalization across diverse agent.
- Feb 28, 2026
- 262K context
- $0.55/M input
- $3.50/M output
- $0.225/M cache read
MoonshotAI: Kimi K2.5TEETrusted Execution Environment
Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed visual and text tokens, it delivers strong performance in general reasoning, visual coding, and agentic tool-calling.
- Jan 29, 2026
- 262K context
- $0.60/M input
- $3.00/M output
- $0.22/M cache read
Qwen: Qwen3 Embedding 8BTEETrusted Execution Environment
The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining.
- Jan 14, 2026
- 33K context
- $0.01/M input
- $0.00/M output
DeepSeek: DeepSeek V3.2TEETrusted Execution Environment
DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios. A scalable reinforcement learning post-training framework further improves reasoning, with reported performance in the GPT-5 class, and the model has demonstrated gold-medal results on the 2025 IMO and IOI. V3.2 also uses a large-scale agentic task synthesis pipeline to better integrate reasoning into tool-use settings, boosting compliance and generalization in interactive environments.
- Dec 3, 2025
- 164K context
- $1.00/M input
- $1.00/M output
- $0.50/M cache read
Qwen: Qwen3 VL 30B A3B InstructTEETrusted Execution Environment
Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception of real-world/synthetic categories, 2D/3D spatial grounding, and long-form visual comprehension, achieving competitive multimodal benchmark results. For agentic use, it handles multi-image multi-turn instructions, video timeline alignments, GUI automation, and visual coding from sketches to debugged UI. Text performance matches flagship Qwen3 models, suiting document AI, OCR, UI assistance, spatial tasks, and agent research.
- Nov 28, 2025
- 128K context
- $0.20/M input
- $0.70/M output
Meta: Llama 3.3 70B InstructTEETrusted Execution Environment
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks.
- Nov 28, 2025
- 131K context
- $2.00/M input
- $2.00/M output
Sentence Transformers: all-MiniLM-L6-v2TEETrusted Execution Environment
The all-MiniLM-L6-v2 embedding model maps sentences and short paragraphs into a 384-dimensional dense vector space, enabling high-quality semantic representations that are ideal for downstream tasks such as information retrieval, clustering, similarity scoring, and text ranking.
- Nov 25, 2025
- 512 context
- $0.005/M input
- $0.00/M output
Qwen2.5 7B InstructTEETrusted Execution Environment
Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2:
- Significantly more knowledge and has greatly improved capabilities in coding and mathematics, thanks to our specialized expert models in these domains.
- Significant improvements in instruction following, generating long texts (over 8K tokens), understanding structured data (e.g, tables), and generating structured outputs especially JSON. More resilient to the diversity of system prompts, enhancing role-play implementation and condition-setting for chatbots.
- Long-context Support up to 128K tokens and can generate up to 8K tokens.
- Multilingual support for over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, Arabic, and more.
- Oct 3, 2025
- 33K context
- $0.10/M input
- $0.20/M output
Google: Gemma 3 27BTEETrusted Execution Environment
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3 27B is Google's latest open source model, successor to
- Oct 3, 2025
- 54K context
- $0.15/M input
- $0.46/M output
- $0.075/M cache read
OpenAI: GPT OSS 120BTEETrusted Execution Environment
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized to run on a single H100 GPU with native MXFP4 quantization. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation.
- Oct 3, 2025
- 131K context
- $0.15/M input
- $0.60/M output
OpenAI: GPT OSS 20BTEETrusted Execution Environment
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference and deployability on consumer or single-GPU hardware. The model is trained in OpenAI’s Harmony response format and supports reasoning level configuration, fine-tuning, and agentic capabilities including function calling, tool use, and structured outputs.
- Oct 3, 2025
- 131K context
- $0.04/M input
- $0.15/M output