Baseten Integrates DeepSeek-V4.1-Flash into Model APIs, Featuring 1M-Token Context – Unite.AI

Introducing DeepSeek-V4.1-Flash: Revolutionizing Model APIs on Baseten

On September 11, 2026, Baseten unveiled the DeepSeek-V4.1-Flash model, a remarkable 552B-parameter multimodal mixture-of-experts (MoE) architecture. This innovative model utilizes 8B active parameters for prefill and 16B for decoding across a vast 1M-token context window, enhancing the capabilities available on their platform. For more insights, you can read Baseten’s official announcement here.

DeepSeek has also made the model’s open weights accessible on Hugging Face, as detailed in their announcement from September 9, 2026. The model supports text and image inputs and generates textual outputs. It’s licensed under the MIT License as per the model card. Baseten describes V4.1-Flash as DeepSeek’s third open-weight release this year, featuring the exclusive Causal Encoder-Decoder design. Future support for Baseten’s Loops training product is on the horizon.

Benchmarking Results: A New Standard for Performance

The model card showcases impressive benchmark results, particularly at the highest reasoning effort setting of 100. V4.1-Flash scored 90.6 on Terminal-Bench 2.1, outperforming V4-Flash (82.7) and V4-Pro (87.9). It scored 74.2 on DeepSWE v1.1, compared to 54.4 and 62.7, and 54.8 on AutomationBench, rising above 37.7 and 43.2 for earlier models. While V4.1-Flash demonstrates superior performance with significantly fewer parameters, it’s important to note that scoring 54.8 on AutomationBench indicates it may struggle with complex workflows, emphasizing the continued need for human oversight in agent operations.

When compared to other leading models, V4.1-Flash registered 90.9 on GPQA Diamond, a Codeforces rating of 3471, and 63.9 on HLE with tools. Notably, this is DeepSeek’s first non-experimental model to handle native image input, a feature previously limited to experimental systems. Its scores of 78.9 on Chartography and 49 on ZeroBench validate its advancements against previous experimental benchmarks.

Innovative Causal Encoder-Decoder Architecture

V4.1-Flash employs a sophisticated 40-layer Transformer configured as a 20-layer causal encoder followed by a 20-layer decoder. The decoder’s global key-value (KV) cache is projected from the final encoder hidden states, enhancing efficiency. With 8B parameters activated during prefill and 16B during decoding, Baseten highlights the model’s cost-effectiveness for coding agents, where prefill tokens significantly outnumber decode tokens.

The model features Compressed Sparse Attention 2, with each layer operating in one of three static modes (Full, Reindex, or Reuse). This design, along with a Hierarchical Sparse Indexer, reduces indexing costs significantly while maintaining performance. The combined innovations cut the global KV cache size to 890 bytes per token—approximately one-quarter of the previous model. Additionally, the SWA Bounded Replay mechanism reconstructs KV states efficiently by only replaying the most recent tokens, reducing the persistent KV footprint to about one-eighth of the earlier generation.

Each MoE layer integrates one shared expert and 384 routed experts, with six experts activated for each token. Notably, the model also introduces Engram conditional memory, hosting 196B parameters alongside DSpark speculative decoding. DeepSeek developed V4.1-Flash from scratch using a 45T-token multimodal corpus, while extending context to 1M tokens after extensive training and fine-tuning processes.

Transitioning to DeepSeek API and Enhanced Service via Baseten

As DeepSeek phases out V4-Flash and V4-Flash-Vision-Exp, the previous API models will temporarily redirect to V4.1-Flash for compatibility. New API pricing took effect on September 10, 2026, with off-peak rates set at 50% of peak rates. Noteworthy partners, including WorkBuddy and OpenCode, fully support V4.1-Flash on their platforms.

Baseten’s Inference Stack efficiently serves this model with NVIDIA Dynamo and KV cache-aware routing, further optimizing request handling. V4.1-Flash will be accessible through Baseten’s Model Library, with dedicated deployments for teams requiring reserved capacity.

Starting September 14, 2026, all deepseek-v4-pro requests will reroute to V4.1-Flash at corresponding rates, a transition expected to improve performance, cost, speed, and overall runtime until V4.1-Pro is released.

Here are five FAQs regarding Baseten’s addition of DeepSeek-V4.1-Flash to Model APIs with a 1M-token context, based on the Unite.AI release:

1. What is DeepSeek-V4.1-Flash?

Answer: DeepSeek-V4.1-Flash is an advanced model integration introduced by Baseten that enhances performance by allowing for a context of up to 1 million tokens. This capability enables users to process and analyze extensive data streams more efficiently, making it particularly useful for applications requiring large datasets.

2. How does the 1M-token context improve model performance?

Answer: The 1M-token context allows the model to retain and analyze significantly more information at once, leading to better understanding and generation of text. This feature is particularly beneficial for tasks that require comprehensive context, such as conversational AI, summarization, or document examination, ultimately resulting in more coherent and relevant outputs.

3. What are the practical applications of using DeepSeek-V4.1-Flash?

Answer: DeepSeek-V4.1-Flash can be applied in various fields, including natural language processing, customer support automation, content generation, and any scenario where deep analysis of large text datasets is necessary. Its ability to handle a 1M-token context means it can support complex projects that require nuanced understanding.

4. What are the benefits of using Baseten’s Model APIs with DeepSeek integration?

Answer: Using Baseten’s Model APIs with DeepSeek integration provides users with a robust toolkit that combines ease of access to advanced AI capabilities with the opportunity to perform complex analytical tasks. The APIs enable seamless integration into existing workflows and applications, facilitating rapid development and deployment of AI solutions.

5. Is there any learning curve associated with implementing DeepSeek-V4.1-Flash?

Answer: While the integration is designed to be user-friendly, some users may need to familiarize themselves with the specifics of the DeepSeek model and its API functionalities. Baseten provides documentation and support to help developers smoothly transition and fully leverage the enhanced capabilities of DeepSeek-V4.1-Flash in their applications.

Source link

Anthropic Integrates Claude Code into Enterprise Offerings

Anthropic Unveils New Subscription Model: Introducing Claude Code for Enterprise

Claude Code Joins the Enterprise Suite

On Wednesday, Anthropic announced an exciting new subscription offering that integrates the popular command-line tool, Claude Code, into its Claude for Enterprise suite. Originally accessible only through individual accounts, this integration allows businesses to leverage sophisticated features and enhanced administrative tools.

A Response to Customer Demand

“This is the most requested feature from our business team and enterprise customers,” said Scott White, Anthropic’s product lead, in an interview with TechCrunch.

Strengthening Competitive Edge

This strategic move places Anthropic in a better position to compete with command-line tools from industry giants like Google and GitHub, which also launched with enterprise-level integrations.

The Rise of Claude Code

Since its launch in June, Claude Code has rapidly gained popularity, offering a unique, agentic approach to command-line programming that sets it apart from traditional IDE-based tools. However, this surge in popularity has led to challenges, particularly for individual users facing unexpected usage limits. The new enterprise subscription addresses these concerns, enabling businesses to establish detailed spending controls that can be adjusted for heavy usage.

Innovative Integrations with Claude.ai

Anthropic is particularly enthusiastic about the potential of integrating Claude Code with the Claude.ai chatbot. This new enterprise model allows for flexible management of both tools, enabling businesses to create Claude Code prompts alongside the chatbot or to incorporate the command-line tool into internal data sources seamlessly.

Transforming Customer Feedback into Action

Scott White highlighted the transformative impact of enterprise integrations involving customer feedback tools. By utilizing Claude to synthesize large volumes of feedback, businesses can translate insights into concrete product improvements. “There’s something magical about blending customer feedback, getting the voice of your customer, and considering solutions you might prototype to meet their unique challenges,” White noted. “It’s something that as a product manager was simply not possible for me even a year ago.”

Join the Conversation

We’re continuously looking to improve! Share your thoughts and feedback on TechCrunch’s coverage and events by filling out our survey. You may even win a prize!

Here are five frequently asked questions (FAQs) regarding Anthropic’s Claude Code and its integration into enterprise plans:

FAQ 1: What is Claude Code?

Answer: Claude Code is an advanced AI language model developed by Anthropic, designed for various tasks such as code generation, debugging, and natural language processing. It aims to assist developers and businesses in automating workflows and enhancing productivity around coding tasks.


FAQ 2: How can businesses benefit from integrating Claude Code into their enterprise plans?

Answer: Businesses can leverage Claude Code to streamline software development processes, improve code quality, and accelerate project timelines. By using AI for repetitive coding tasks, teams can focus on higher-level design and problem-solving, leading to increased efficiency and innovation.


FAQ 3: What are the key features of Claude Code in the enterprise plans?

Answer: Key features include code generation capabilities, bug detection and fixing, natural language queries to code, integration with existing development tools, and continuous learning from user interactions to improve performance over time. Additionally, enterprise plans may offer enhanced security, scalability, and support.


FAQ 4: Is Claude Code customizable for specific business needs?

Answer: Yes, Claude Code can be tailored to fit specific business requirements, allowing integration with existing workflows, frameworks, and coding languages. Customization options may include training the model on proprietary data to enhance relevance and accuracy.


FAQ 5: How is the pricing structured for enterprise plans that include Claude Code?

Answer: Pricing for enterprise plans incorporating Claude Code typically depends on factors such as user count, usage volume, and specific features selected. Custom quotes can be provided based on individual business needs, ensuring flexibility and scalability as companies grow. For detailed pricing information, it’s best to contact Anthropic directly.


Feel free to adjust any of the FAQs or answers based on your specific needs or context!

Source link