Tag: AI Models

  • Google Launches Gemini 3.8 Live and Extended Thinking: Production-Ready Voice AI at a Fraction of the Cost

    Google Launches Gemini 3.8 Live and Extended Thinking: Production-Ready Voice AI at a Fraction of the Cost

    Google DeepMind released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026, a pair of real-time audio models designed to power production-grade voice agents. The launch places Google at the top of independent quality benchmarks while offering pricing that undercuts rival frontier models by more than 50 percent. For developers and enterprises building voice applications, the announcement marks a meaningful shift in what is accessible at scale.

    What Was Announced

    Google DeepMind introduced two distinct models on September 15, 2026. Gemini 3.8 Live is optimized for speed and cost efficiency, targeting high-volume deployments such as customer support, scheduling, and tutoring applications. Gemini 3.8 Live Extended Thinking is the higher-capability variant, built for complex agentic tasks that require the model to reason carefully before responding.

    Both models are available immediately through the Gemini API and Google AI Studio. They are also integrated across Google’s own products, including Gemini Enterprise, Google Workspace, Search Live, and the consumer Gemini Live app. This broad rollout positions the models not just as developer tools but as infrastructure embedded in services used by hundreds of millions of people daily.

    The announcement arrives less than two weeks after OpenAI opened GPT-Live-1, its competing full-duplex voice model, to developers at $0.05 per minute. Google’s move signals an escalating race to dominate the production voice agent market, a segment seen as one of the highest-growth areas in enterprise AI adoption.

    A key differentiator is Extended Thinking, a mode that allows the model to reason through difficult queries, use external tools, and retrieve information before speaking, all while keeping the conversation feeling natural and uninterrupted. Google says this addresses a persistent criticism of voice AI: that capable models pause too long or produce unnatural turn-taking when asked to think.

    Technical Details

    Gemini 3.8 Live processes audio natively, without transcribing speech to text and then back to speech again. This end-to-end approach preserves prosody, reduces latency, and lets the model pick up on tone and speaking pace as contextual signals. The result is conversation behavior that responds to how someone speaks, not just what they say.

    Gemini 3.8 Live Extended Thinking introduces a reasoning layer that activates on demand for complex queries. This enables the model to invoke tools, query external APIs, and reason over documents without surfacing that computational work to the caller. Developers control reasoning depth via a thinking budget API parameter, allowing them to trade off latency against task complexity at the application level.

    On Artificial Analysis’ Speech to Speech Quality Index, Gemini 3.8 Live Extended Thinking scored 82.6, the highest overall score recorded on the benchmark. It also leads in agentic task completion with a score of 68.6 percent, outperforming all other models tested. Pricing is set at $0.005 per minute for audio input and $0.018 per minute for audio output, which translates to approximately $0.84 per hour for the standard model on Artificial Analysis’ cost-per-hour measure. The Extended Thinking variant costs $3.50 per hour on the same measure.

    The models integrate with Google’s existing infrastructure tools, including function calling, code execution, and grounding with Google Search. These capabilities were previously available in Gemini’s text-based API but are now surfaced natively in a voice context, letting developers build voice agents that search, calculate, and execute without switching modalities.

    Industry Impact and Reactions

    The pricing structure is a central part of the story. OpenAI’s GPT-Live-1 is billed at $0.05 per minute, which translates to roughly $3 per hour for voice input alone, before adding the cost of the underlying reasoning model. Google’s $0.84 per hour for Gemini 3.8 Live undercuts that figure by more than 70 percent. Even the more capable Extended Thinking variant at $3.50 per hour is competitive at the top of the market.

    For enterprise buyers evaluating build-versus-buy decisions on voice pipelines, cost at scale is a primary factor. The differential gives Google an opening to win deployments where conversation quality at the standard tier is sufficient and where budget constraints have previously ruled out frontier-quality voice AI. Call center automation, appointment scheduling, and tutoring platforms are all cited as target use cases.

    The release also adds competitive pressure to Eleven Labs, Deepgram, and other specialized voice AI providers. These companies have built market position on low-latency, high-quality text-to-speech and speech-to-text tooling. A general-purpose voice reasoning model from a hyperscaler, priced below most point solutions and integrated directly into Google Workspace, changes the calculus for many buyers. Developer reaction was broadly positive, with particular attention on the Extended Thinking variant’s benchmark performance and the elimination of the awkward-pause problem through the thinking budget mechanism.

    What Comes Next

    Google has not announced a specific date for the next set of Gemini 3.8 Live features, but the company indicated at launch that multimodal input, specifically the ability to process live video alongside audio, is on the near-term roadmap. This would extend the models’ utility beyond phone-style voice agents into video call copilots and real-time translation applications.

    Current pricing is locked through at least January 1, 2027, when Google has stated that token rates for several Gemini 3.8 models will approximately double. Developers building on the current pricing window have roughly three and a half months to evaluate production workloads before a rate adjustment. Google’s track record of extending promotional pricing windows suggests the transition may be gradual, but enterprise customers are advised to model both scenarios.

    Conclusion

    Google’s Gemini 3.8 Live launch combines benchmark-leading performance with pricing that meaningfully expands the market for production voice AI. Whether the goal is a customer support agent, a scheduling assistant, or a more capable consumer application, the two new models offer developers a credible new option that trades on both quality and cost. As voice becomes an increasingly central interface for AI products, the race to own that layer is accelerating, and Google has moved to the front of the pack on the metrics that matter most.

    Stay updated on the latest AI news at Evolve Digital.

  • Moonshot AI Releases Kimi K3: The World’s Largest Open-Weight AI Model at 2.8 Trillion Parameters

    Moonshot AI Releases Kimi K3: The World’s Largest Open-Weight AI Model at 2.8 Trillion Parameters

    On July 16, 2026, China’s Moonshot AI released Kimi K3, a 2.8-trillion-parameter model that instantly became the largest open-weight AI release in history. The model surpasses every previous open-weight system by a wide margin and arrives at a moment when Chinese AI labs are demonstrating an ability to match or approach U.S. frontier systems despite significant restrictions on advanced chip exports. Kimi K3 is available via API today and Moonshot AI has committed to releasing full open weights by July 27, 2026.

    What Was Announced

    Moonshot AI, the Beijing-based startup behind the Kimi series of AI products, launched Kimi K3 via its website and API on July 16, 2026. The company describes it as “the world’s first open 3T-class model” — shorthand for a model in the 3-trillion-parameter class — and the release has already drawn attention from major technology outlets including Bloomberg, VentureBeat, and Tom’s Hardware.

    The launch is significant not only for its technical scale but for its timing. Kimi K3 arrives days after Google’s Gemini 3.5 Pro debuted on July 17 and less than two weeks after OpenAI broadly released GPT-5.6. The result is one of the most competitive weeks in AI development history, with a Chinese open-weight model sitting alongside the latest closed U.S. frontier systems on benchmark leaderboards.

    Moonshot AI has promised to release the model’s full weights publicly by July 27, 2026, placing it under an open license for developers worldwide. As of the API launch date, Kimi K3 is accessible at $0.30 per million cache-hit input tokens, $3 per million on cache misses, and $15 per million output tokens.

    In its own benchmark reporting, Moonshot places Kimi K3 ahead of Claude Opus 4.8 and GPT-5.5, with only Claude Fable 5 and GPT-5.6 Sol ranking higher across most tasks evaluated. Independent third-party evaluations on coding benchmarks, including the Frontend Code Arena, have shown similar results.

    Technical Details

    Kimi K3 uses a Mixture-of-Experts (MoE) architecture with 896 expert sub-networks. For any given input token, the model activates just 16 of those experts — roughly 1.8 percent of the total pool — meaning the effective compute per forward pass corresponds to approximately 41 billion active parameters, rather than the full 2.8 trillion. This design allows the model to pack enormous capacity into its weights while keeping inference costs at a level competitive with much smaller dense models.

    The model was trained on 45 trillion tokens of multimodal data spanning text, images, audio, and video, giving it native reasoning ability across all four content types. Its context window extends to 1 million tokens, designed specifically for long-horizon tasks such as processing large codebases, extended documents, or complex multi-step agent workflows.

    Moonshot built Kimi K3 with compute efficiency as a priority constraint, given U.S. export controls that have limited Chinese labs’ access to the most advanced Nvidia chips. The architecture choices — sparse expert activation, efficient attention mechanisms for long context, and a large total parameter count relative to active compute — reflect an engineering approach optimized to extract maximum capability from available hardware.

    Industry Impact and Reactions

    The Kimi K3 release is another data point in a clear trend: Chinese AI laboratories are closing the gap with U.S. frontier systems faster than most industry observers predicted, and they are doing so while operating under chip restrictions that were expected to slow their progress significantly. Kimi K3’s self-reported performance, showing it outperforming models that cost far more to serve, demonstrates that parameter efficiency and scale can partially offset the compute disadvantage.

    For the open-source and open-weight AI community, the release is particularly notable. The largest open-weight models available before Kimi K3 sat well below one trillion parameters. A 2.8-trillion-parameter system with promised downloadable weights fundamentally changes what researchers, enterprises, and developers working outside of major cloud providers can access and fine-tune. The Apache License under which the model is expected to be released adds further flexibility for commercial use.

    The competitive context matters for U.S. frontier labs as well. OpenAI, Anthropic, and Google now face a public benchmark comparison from an open model that competes seriously on coding and multimodal reasoning tasks — and that any organization can download, run privately, and modify. This shifts the calculus for enterprises evaluating proprietary versus open systems, particularly those with data privacy or sovereignty requirements that make cloud-only deployments difficult.

    What Comes Next

    The most anticipated near-term milestone is the open-weights release Moonshot AI has committed to by July 27, 2026. Once the full model checkpoints are available on Hugging Face, independent researchers and benchmark organizations will be able to conduct thorough third-party evaluations, which may confirm, revise, or challenge the self-reported numbers Moonshot published at launch. Early community reception of the API has been positive on coding and agent benchmarks.

    Moonshot AI has also positioned Kimi K3 as a foundation for its enterprise customization ecosystem. Developers who want to use the model as a starting point for fine-tuned, task-specific deployments can do so once the weights are public. This mirrors the approach taken by Meta with the Llama series, and it suggests that Moonshot is competing not just on raw model performance but on building an open AI ecosystem anchored around a flagship model.

    Conclusion

    Kimi K3 marks a genuine inflection point for open-weight AI development. With 2.8 trillion parameters, a 1-million-token context window, and benchmark results that rival closed frontier models from OpenAI and Anthropic, it resets expectations for what open models can deliver. Its imminent full release will place this capability directly in the hands of developers and researchers globally, at a moment when access to high-performing, customizable AI has rarely mattered more. Moonshot AI’s release confirms that the frontier of AI development is no longer confined to a handful of U.S. laboratories.

    Stay updated on the latest AI news at Evolve Digital.

  • Google Scraps and Rebuilds Gemini 3.5 Pro Ahead of July 17 Launch: What We Know

    Google Scraps and Rebuilds Gemini 3.5 Pro Ahead of July 17 Launch: What We Know

    In a significant departure from standard AI development practice, Google disclosed on July 16, 2026 that it completely scrapped and rebuilt the base model for Gemini 3.5 Pro after critical structural failures emerged during enterprise testing on Vertex AI. The original architecture exhibited performance gaps across three core capabilities that Google engineers deemed unacceptable for a product competing at the frontier of AI development. Rather than attempting to patch the existing model through fine-tuning, Google DeepMind chose a full pre-training rebuild from scratch. The rebuilt Gemini 3.5 Pro is now targeting a launch on July 17, 2026, though Google has not officially confirmed the date, pricing, or technical specifications as of this writing.

    What Was Announced

    Google’s decision to restart Gemini 3.5 Pro’s development from the ground up came after enterprise testing on Vertex AI revealed failures across three critical capability categories. Engineers identified recursive tool-calling instability, which is a fundamental requirement for agentic coding workflows that businesses rely on to automate complex software development tasks. The original model also struggled with complex SVG scene generation, failing to reliably produce accurate vector graphics output. A third category of failure involved mathematical reasoning, where the model showed performance gaps compared to what Google considered acceptable for a flagship product.

    The issues were described as structural rather than addressable through standard post-training techniques such as fine-tuning or reinforcement learning. This distinction is significant: fine-tuning can improve model behavior within the constraints of an existing architecture, but structural failures require rebuilding the foundation. Google made the call to conduct a new pre-training cycle rather than ship a model with foundational weaknesses.

    The rebuilt model reportedly addresses these shortcomings with a new focus on front-end generation capabilities. Reported improvements include greater precision in UI design generation, more concise and reliable code output, improved 3D modeling performance, and stable multi-step agent tool-calling. These capabilities target the enterprise and developer markets where Gemini 3.5 Pro will compete most directly.

    Pricing reported for the model is approximately $15 per million input tokens and $60 per million output tokens, though Google has not officially confirmed these figures. Access to the Deep Think reasoning tier, which enables more extended chain-of-thought reasoning, is expected to be gated behind the $250/month Gemini Ultra subscription.

    Technical Details

    Among the most significant reported specifications is a 2 million token context window, which would represent a substantial lead over competing models. Most frontier models currently support context windows in the range of 1 million tokens. A 2 million token context would allow developers to process entire large codebases, comprehensive legal documents, or extended research archives in a single inference call, enabling new categories of enterprise workflows that are currently impractical with smaller context limits.

    The Deep Think reasoning layer is designed to operate as a tiered capability, engaging extended multi-step reasoning for complex tasks while maintaining standard inference speed for simpler requests. This approach mirrors similar reasoning tiers offered by competing models, including extended thinking modes in Anthropic’s Claude family and OpenAI’s reasoning model lineup. The practical effect is that developers can route simpler queries to standard inference and reserve Deep Think for tasks that require sustained logical chains.

    What has not been confirmed officially includes the model’s parameter count, the specific training data composition, infrastructure details, and full benchmark performance across standard evaluation suites. Until Google publishes an official model card and benchmark results, all technical specifications should be treated as reported rather than verified.

    Industry Impact and Reactions

    The Gemini 3.5 Pro rebuild places Google in direct competition with recently released frontier models that have set new performance benchmarks. Anthropic’s Claude Fable 5 has posted leading scores on SWE-bench Pro, a widely used software engineering benchmark, which observers have flagged as the current bar for agentic coding capability. OpenAI’s GPT-5.6 Sol, released earlier in July 2026, has similarly established strong positions in coding, scientific reasoning, and knowledge work. Google’s decision to delay rather than ship an architecturally flawed model signals that it is treating Gemini 3.5 Pro as a competitive flagship, not a routine product update.

    The pricing structure, if confirmed, positions Gemini 3.5 Pro in the premium tier of frontier model pricing. At approximately $15 per million input tokens and $60 per million output tokens, it sits above efficiency-focused tiers but within the range of models targeting demanding enterprise use cases. The Deep Think tier’s inclusion in the $250/month Ultra subscription rather than per-token pricing represents a bet on subscription adoption among enterprise customers who want predictable costs for complex reasoning workloads.

    Google simultaneously plans to launch Nano Banana Pro, a separate image generation model targeting competition with OpenAI’s GPT-Image 2. This dual-launch strategy suggests Google is attempting to address both language model and image generation markets simultaneously, potentially to capture developer attention ahead of competing model releases expected later in Q3 2026. The combination of a rebuilt language model and a new image model would represent Google’s most comprehensive AI product push since the original Gemini launch.

    What Comes Next

    The reported launch date of July 17, 2026 means developers and enterprises should watch for official API availability, model card publication, and benchmark disclosure within the next 24 hours. Google has not officially confirmed the date as of July 16, so any slippage remains possible given the scale of the architectural rebuild. When benchmarks do arrive, the comparisons that will matter most are performance on SWE-bench Pro for agentic coding capability and MMLU for general reasoning, where the rebuilt model’s results will clarify whether the full pre-training cycle achieved its intended improvements.

    Longer term, the launch will provide the first concrete data point on whether Google’s willingness to absorb a development delay translates into the kind of architectural quality that developers and enterprise customers reward with adoption. The competitive window is narrow: with Anthropic and OpenAI both releasing models on faster cadences, Google will need Gemini 3.5 Pro to establish a clear performance or capability differentiation to hold its position in the enterprise AI market.

    Conclusion

    Google’s decision to scrap and rebuild Gemini 3.5 Pro reflects a broader maturation in how frontier AI labs approach model quality under competitive pressure. The willingness to accept a delayed release rather than ship a model with structural weaknesses in tool-calling, SVG generation, and mathematical reasoning signals that architectural integrity is becoming as important as release cadence in the competition for enterprise AI adoption. As the model prepares for its reported July 17 launch, the industry will be watching closely to see whether the rebuild delivers on the performance improvements Google DeepMind targeted, and whether a 2 million token context window proves to be the differentiator Google needs.

    Stay updated on the latest AI news at Evolve Digital.

  • OpenAI Releases GPT-5.6 Sol, Terra, and Luna: Three Frontier Models Go Public After Government Security Review

    OpenAI Releases GPT-5.6 Sol, Terra, and Luna: Three Frontier Models Go Public After Government Security Review

    OpenAI made its most significant model release of 2026 on July 9, launching three new GPT-5.6 models to the public simultaneously: Sol, Terra, and Luna. The rollout came after a 12-day delay requested by the US government over national security concerns, marking the first time a major AI model release was formally held pending a White House security evaluation. All three models are now available to ChatGPT subscribers and API developers worldwide, representing a major expansion of OpenAI’s publicly accessible frontier AI offerings.

    What Was Announced

    OpenAI released GPT-5.6 as a family of three distinct models rather than a single flagship, each positioned to serve a different tier of user and use case. Sol is the top-tier variant optimized for frontier reasoning and long-horizon agentic work, priced at $5 per million input tokens and $30 per million output tokens. Terra is a balanced, everyday model designed to match or exceed GPT-5.5 performance at approximately half the cost, priced at $2.50 per million input tokens and $15 per million output tokens. Luna is the fastest and most affordable option in the family at $1 per million input tokens and $6 per million output tokens.

    The announcement was anticipated for several days before the July 9 launch date was confirmed. OpenAI had originally planned an earlier release but agreed to a delay after the US government raised national security concerns about potential misuse. After a 12-day evaluation process involving White House officials, OpenAI received clearance to proceed with a global rollout.

    All three models are now accessible via the ChatGPT interface and OpenAI’s API. GPT-5.6 Sol targets developers and enterprises building complex agentic pipelines, while Terra and Luna serve broader audiences including standard ChatGPT subscribers on various plan tiers.

    The three-model structure echoes how OpenAI has tiered previous releases, but the inclusion of a government security review as a formal pre-release checkpoint represents a new pattern for the company and potentially for the industry at large.

    Technical Details

    GPT-5.6 Sol is built for long-horizon agentic work, a class of tasks that require a model to plan and execute multi-step processes over extended periods. The model introduces a new max reasoning effort setting, which allows developers to instruct the model to apply deeper reasoning passes to problems that benefit from extended computation. Sol also features an ultra mode, designed for faster completion of complex tasks without sacrificing the model’s reasoning depth.

    Terra is positioned as the everyday workhorse of the GPT-5.6 family. OpenAI describes Terra as delivering GPT-5.5-competitive performance at roughly 2x lower cost, making it an economically practical choice for organizations running large volumes of inference at near-frontier capability levels. Luna targets the high-throughput end of the market, prioritizing speed and cost efficiency over raw reasoning depth.

    The full-duplex voice capability introduced earlier this week with GPT-Live is not directly part of the GPT-5.6 release, but GPT-Live delegates complex queries to frontier models in the background. With GPT-5.6 now publicly available, future updates to the voice product may incorporate the new model family as the underlying reasoning backbone for those delegated tasks.

    Industry Impact and Reactions

    The July 9 launch places OpenAI back at the frontier of publicly available commercial AI after a period marked by export control disruptions and model delays. The simultaneous availability of Sol, Terra, and Luna across the API gives developers immediate access to a tiered set of frontier options, a contrast to the phased rollouts that characterized some prior OpenAI releases.

    The pricing structure is noteworthy in the current competitive landscape. Terra at $2.50 per million input tokens directly competes with Anthropic’s Claude Sonnet 5, which is available at $2 per million input tokens through August 31 at introductory pricing. Luna at $1 per million input tokens positions OpenAI competitively in the high-volume, cost-sensitive segment of the market where speed and price are the primary purchasing criteria.

    The government review process that preceded this launch is a notable development for the industry as a whole. AI companies have faced increasing pressure from legislators and national security officials to provide advance notice and allow evaluation of their most capable models before public release. The 12-day White House evaluation of GPT-5.6 suggests this informal framework may be becoming a de facto step in the release pipeline for frontier AI systems.

    What Comes Next

    Speculation about GPT-6 has intensified in recent weeks, with several industry analysts suggesting an announcement could come before the end of 2026. The rapid succession of GPT-5.5, GPT-Live, and now GPT-5.6 within a compressed window suggests OpenAI is accelerating its release cadence as competitive pressure mounts from Anthropic, Google DeepMind, and international AI developers. OpenAI has not confirmed a GPT-6 timeline.

    For enterprise and developer customers, the immediate priority will be evaluating where each GPT-5.6 variant fits their existing workflows. Organizations that built pipelines around GPT-5.5 will need to benchmark Terra and Sol against their current performance baselines before migrating. OpenAI has indicated that GPT-5.5 will remain available in the API for the near term, giving developers time to assess the new family at their own pace.

    Conclusion

    OpenAI’s release of GPT-5.6 Sol, Terra, and Luna on July 9, 2026 expands the frontier of publicly available AI with a three-tier model family covering agentic reasoning, balanced everyday performance, and high-speed cost-efficient inference. The unusual inclusion of a government security review before launch marks a shift in how regulators and AI companies are managing the release of the most capable models. With pricing that directly competes across multiple market segments, the GPT-5.6 family arrives as one of the more consequential OpenAI releases of the year.

    Stay updated on the latest AI news at Evolve Digital.

  • Anthropic Releases Claude Opus 4.8 With Dynamic Workflows and Major Coding Improvements

    Anthropic Releases Claude Opus 4.8 With Dynamic Workflows and Major Coding Improvements

    Anthropic has released Claude Opus 4.8, the latest iteration of its flagship AI model, bringing meaningful gains in coding reliability, reasoning, and autonomous operation. Released on May 29, 2026, just 41 days after Opus 4.7, the update introduces a headline new capability called Dynamic Workflows and delivers measurable benchmark improvements across core performance areas. The model is available globally today via the Anthropic API and Claude.ai at the same price point as its predecessor.

    What Was Announced

    Anthropic described Claude Opus 4.8 as offering “sharper judgment, more honesty about its progress, and the ability to work independently for longer than its predecessors.” The company released benchmark data showing improvements on two key metrics: agentic coding performance rose from 64.3% to 69.2%, while multidisciplinary reasoning with tools improved from 54.7% to 57.9%.

    One of the more notable reliability improvements is in code quality oversight. Anthropic says Opus 4.8 is approximately four times less likely than Opus 4.7 to allow flaws in code it has written to pass silently without flagging them, addressing a persistent pain point for teams relying on AI models in software development pipelines.

    Speed also improved: the Opus 4.8 fast mode is roughly 2.5 times quicker than the equivalent mode in Opus 4.7. Critically, Anthropic kept pricing identical to the previous model version, meaning existing API users receive the full upgrade at no additional cost.

    The centerpiece of the release is Dynamic Workflows, now available in research preview. This feature is designed to enable Opus 4.8 to coordinate and manage complex, long-horizon tasks by orchestrating hundreds of parallel subagents simultaneously. Anthropic positioned this capability specifically for enterprise teams building large-scale agentic pipelines where multiple AI instances must collaborate on a shared goal.

    Technical Details

    Dynamic Workflows represents a significant architectural extension of how Claude operates in multi-agent contexts. Rather than functioning as a single model responding sequentially, Opus 4.8 with Dynamic Workflows acts as an orchestrator, delegating subtasks to parallel subagents and synthesizing their outputs into coherent results. This allows the model to tackle problems that would be impractical to complete within a single context window or within the latency constraints of a linear workflow.

    The coding improvements in Opus 4.8 are tied closely to enhancements in self-monitoring. The model shows improved ability to recognize when its own output contains errors or uncertainties, and to flag these rather than proceeding with flawed assumptions. This behavioral shift is particularly significant in autonomous coding scenarios, where silent errors can propagate through large codebases before being detected.

    Anthropic also notes that fast mode throughput improvements were achieved through inference optimizations rather than model compression, preserving the underlying capability profile of the model while significantly reducing latency for time-sensitive applications.

    Industry Impact and Reactions

    The release comes in a period of rapid iteration across the frontier AI model landscape. Anthropic’s 41-day release cycle from Opus 4.7 to 4.8 signals a faster cadence than the company has historically maintained, reflecting competitive pressure from OpenAI and Google, both of which have accelerated their own release timelines in 2026.

    The combination of Dynamic Workflows and improved coding reliability is directly relevant to the growing enterprise market for agentic AI. Businesses deploying AI in software development, data analysis, and automated workflow management stand to benefit most from the improvements. The fact that the upgrade carries no price increase removes one of the traditional adoption barriers for enterprise customers already on the Anthropic API.

    Claude Opus 4.8 also arrives alongside a significant financial milestone for Anthropic: the company recently raised additional private funding, reaching a valuation of approximately $965 billion. This financial backdrop gives Anthropic substantial runway to continue research investment and infrastructure expansion as it competes at the frontier of large language model development.

    What Comes Next

    Dynamic Workflows is currently in research preview, suggesting Anthropic is gathering feedback before a broader production release. The company has not announced a specific general availability date for the feature, but the research preview designation typically precedes a full rollout within weeks to months. Anthropic is also expected to bring its next class of models, which the company has referred to informally as Mythos-class, to a wider set of customers later in 2026.

    For teams already using Opus 4.7, the path to Opus 4.8 requires only updating to the latest model version in the API — no integration changes are needed to access the core improvements. Teams interested in Dynamic Workflows will need to apply for the research preview through Anthropic’s developer portal.

    Conclusion

    Claude Opus 4.8 represents a focused, evidence-based upgrade to one of the leading frontier AI models currently available. With improved coding reliability, faster inference, and the introduction of Dynamic Workflows, Anthropic is addressing the real-world needs of developers and enterprises building agentic AI systems. The decision to maintain existing pricing makes this a straightforward upgrade for current users, and positions Anthropic competitively as the race to deploy capable, reliable AI agents in enterprise environments continues to intensify.

    Stay updated on the latest AI news at Evolve Digital.

  • Nvidia CEO Jensen Huang Unveils Ising: The World First Family of Open-Source Quantum AI Models

    Nvidia CEO Jensen Huang Unveils Ising: The World First Family of Open-Source Quantum AI Models

    Nvidia CEO Jensen Huang announced the creation of Nvidia Ising, described as the world first family of open-source quantum AI models, on May 9, 2026. The announcement positions Nvidia at the intersection of two of the most consequential technology bets of the decade: large-scale AI and quantum computing. While commercially viable quantum computing remains years away, the Ising model family represents Nvidia opening move in defining what AI-optimized quantum software might look like when that hardware becomes available.

    What Was Announced

    Jensen Huang announced at an investor event that Nvidia had developed the Ising model family, a set of open-source AI models designed to interface with and accelerate optimization problems that quantum computing architectures are particularly suited to solve. The name references the Ising model from statistical mechanics, a mathematical framework used to model spin interactions in physical systems that has become a foundational benchmark problem for quantum computers.

    The models are being released as open source, consistent with Nvidia strategy across several of its AI research initiatives. Making the models publicly available allows the broader quantum computing and AI research communities to build on them, accelerating development of the tools and workflows needed to make quantum-classical hybrid computing practical for real workloads. Nvidia has positioned itself not as a quantum hardware company but as a software and systems integrator that can bridge quantum hardware from companies like IonQ, IBM, and others with the AI frameworks that developers already know.

    Nvidia described Ising as part of its broader push to integrate quantum computing into its simulation and optimization workflows. The company has existing quantum computing partnerships and has incorporated quantum circuit simulation into its cuQuantum software library. Ising extends that foundation toward AI-native interfaces for quantum problem-solving.

    Technical Details

    The Ising model family is designed around optimization problems — a class of computations that quantum hardware handles particularly well compared to classical systems. Optimization problems appear throughout AI and industrial applications: scheduling, logistics, financial portfolio construction, drug molecule discovery, and materials science simulations are all domains where quantum-optimized solutions could offer significant advantages when hardware matures.

    The models are designed as open-source artifacts that developers can adapt to specific problem domains. Nvidia approach of releasing them under an open license means the research community can extend them to new problem types and hardware backends without waiting for proprietary tools. This positions Nvidia standards and frameworks as the natural foundation for quantum AI development even before quantum hardware achieves commercial viability.

    Nvidia already operates one of the most widely adopted AI software stacks through CUDA, cuDNN, and its associated ecosystem. Extending that stack into the quantum domain through open-source models follows the same playbook: establish the software foundation early and let hardware adoption follow. When commercial quantum hardware eventually arrives at meaningful scale, developers trained on Nvidia quantum tools will likely continue using them.

    Industry Impact and Reactions

    The announcement has drawn attention from both the AI and quantum computing communities. For quantum computing researchers, Nvidia entry as an open-source model provider lends significant institutional weight to efforts to define quantum AI standards. For AI developers, the announcement signals that the GPU giant is thinking seriously about what comes after classical accelerators, even if the timeline remains uncertain.

    Nvidia is not the first major technology company to invest in quantum AI research. Google, IBM, and Microsoft have all built significant quantum computing programs, and all have explored the intersection of quantum hardware with AI workloads. But Nvidia unique position as the dominant supplier of AI training and inference infrastructure gives its quantum AI efforts a distinctive reach: when Nvidia defines what quantum AI software looks like, developers who depend on CUDA have strong incentives to align with that vision.

    Financial analysts covering Nvidia noted that the Ising announcement does not affect the company near-term revenue outlook, which remains overwhelmingly dependent on classical GPU sales. But for investors with a multi-decade horizon, the move is consistent with a pattern of early positioning in transformative technology categories that Nvidia has executed successfully across GPU computing, deep learning, and autonomous vehicles.

    What Comes Next

    Nvidia has not disclosed a specific timeline for when Ising models will be available for download or what quantum hardware backends will be supported at launch. The company is expected to share additional technical details at a forthcoming developer event. In the meantime, the announcement is likely to drive collaboration between Nvidia and quantum hardware providers eager to align their roadmaps with Nvidia open-source software infrastructure.

    Broader commercial quantum advantage in optimization problems is generally expected to emerge in the early-to-mid 2030s based on current hardware trajectories. The Ising model release positions Nvidia to be the software ecosystem of choice when that transition happens.

    Conclusion

    Nvidia release of the Ising open-source quantum AI model family is an early but strategically significant move in what may become one of the most important technology transitions of the coming decade. By establishing an open-source software foundation at the intersection of AI and quantum computing now, Nvidia is following the same playbook that made it the dominant force in classical AI infrastructure — planting a flag early, building developer alignment, and waiting for hardware to mature around its software ecosystem.

    Stay updated on the latest AI news at Evolve Digital.

  • OpenAI Releases GPT-5.5 Instant as ChatGPT New Default Model, Cutting Hallucinations by 52 Percent

    OpenAI Releases GPT-5.5 Instant as ChatGPT New Default Model, Cutting Hallucinations by 52 Percent

    OpenAI rolled out GPT-5.5 Instant as the new default model powering ChatGPT on May 5, 2026, replacing GPT-5.3 Instant and marking the latest step in the company rapid iteration on its flagship conversational AI. The update delivers a significant reduction in hallucinated claims, with OpenAI reporting that GPT-5.5 Instant produces 52.5% fewer hallucinated facts than its predecessor on high-stakes prompts covering medicine, law, and finance. The model is also rolling out as the chat-latest option in the API, meaning developers who have not pinned to a specific model version will automatically receive the upgrade.

    What Was Announced

    OpenAI confirmed on May 5, 2026, that GPT-5.5 Instant would replace GPT-5.3 Instant as the default model in ChatGPT across its web and mobile interfaces. The rollout affects all subscription tiers, making GPT-5.5 Instant the model that free users, Plus subscribers, Pro subscribers, and enterprise customers all encounter by default. API customers using the chat-latest endpoint also receive the upgrade automatically.

    The headline performance improvement is a 52.5% reduction in hallucinated claims on high-stakes prompts. OpenAI defines hallucinated claims as factually incorrect statements presented with apparent confidence, and specifically measured the improvement in domains where accuracy carries significant consequences: medical information, legal analysis, and financial guidance. These are areas where ChatGPT is increasingly used in professional contexts, and where confident errors can cause real harm.

    The update also includes enhanced personalization capabilities, leveraging memory from past conversations, uploaded files, and for users who have connected their Gmail accounts, context from their email. This personalization feature is rolling out to Plus and Pro users on the web first, with mobile support and expansion to additional subscription tiers to follow in the coming weeks.

    Technical Details

    The 52.5% hallucination reduction reflects improvements across several training dimensions. OpenAI has consistently improved factual accuracy through a combination of better training data curation, expanded use of reinforcement learning from human feedback (RLHF), and techniques that train models to self-check outputs before finalizing responses. The specific improvements in medical, legal, and financial domains suggest targeted work on those knowledge areas during fine-tuning.

    GPT-5.5 Instant is positioned as an efficiency-optimized model for fast inference and broad deployment rather than maximum capability on complex reasoning tasks. It sits alongside GPT-5.5 full and reasoning-specialized models like o3 and o4 in the OpenAI lineup. The Instant variant is tuned specifically for the latency requirements of a conversational product used by hundreds of millions of people daily.

    The personalization features represent a shift toward more proactive context ingestion. Earlier memory capabilities required users to explicitly tell the model to remember things. The new approach ingests context from past sessions, files, and connected accounts more automatically, allowing the model to surface relevant information without being prompted.

    Industry Impact and Reactions

    The release comes as OpenAI faces intensifying competition from Anthropic Claude, Google Gemini, and a growing roster of open-weight model providers. The hallucination reduction metric is particularly targeted at enterprise customers, many of whom cite factual reliability as their primary concern about deploying AI in high-stakes workflows. A 52.5% improvement on that dimension is a meaningful competitive differentiator if it holds in independent evaluation.

    The tiered model strategy, with Instant variants optimized for speed, full versions for general capability, and reasoning models for complex tasks, mirrors what both Anthropic and Google have deployed. The AI industry appears to have converged on multi-model architectures as the standard approach for commercial deployment at scale.

    What Comes Next

    OpenAI has indicated that enhanced personalization features will expand to additional data sources and subscription tiers. ChatGPT Go is now available in eight additional European countries and is also being updated to run on GPT-5.5 Instant. The next major version of the GPT-5.5 series is expected to follow OpenAI ongoing release cadence.

    Conclusion

    The release of GPT-5.5 Instant as ChatGPT new default represents meaningful progress on one of the most persistent criticisms of AI language models: the tendency to present inaccurate information with confidence. The 52.5% hallucination reduction is a number that enterprise buyers will notice, and the deeper personalization features reflect OpenAI push to make ChatGPT indispensable in users daily workflows.

    Stay updated on the latest AI news at Evolve Digital.

  • Anthropic’s Secret ‘Mythos’ AI Model Exposed in Data Leak, Described as Step-Change in Capability

    Anthropic’s Secret ‘Mythos’ AI Model Exposed in Data Leak, Described as Step-Change in Capability

    Anthropic is developing a powerful new AI model internally codenamed “Mythos,” according to details that emerged from an accidental data exposure in late March 2026. The leak, first reported by Fortune, revealed that Anthropic considers Mythos its most capable model to date — a significant step up from the Claude 4 family — and has flagged unprecedented cybersecurity concerns associated with its development. The revelation offers a rare window into the advanced frontier work happening inside one of the AI industry’s most safety-conscious labs.

    What Was Revealed

    The existence of Mythos came to light through an inadvertent exposure of internal data, the specifics of which Anthropic has not fully disclosed. In a statement confirming the model’s existence, Anthropic described Mythos as representing a “step change” in capabilities compared to its current production models. The company stopped short of providing a release timeline, benchmark scores, or detailed architectural information, but the internal framing — calling it the most powerful model the company has built — signals an ambitious leap beyond Claude Opus 4.6.

    Anthropic simultaneously disclosed that the development of Mythos has raised internal cybersecurity concerns of an unprecedented nature. The company characterized these concerns as distinct from standard model safety evaluations, suggesting the lab may be grappling with new categories of risk that arise when models reach higher capability thresholds. No specifics were shared about the nature of the threats identified.

    Sources familiar with the situation told Fortune that Mythos is natively multimodal and has demonstrated reasoning and autonomous task completion abilities that substantially exceed those of Claude Opus 4.6 in internal testing. The model’s name evokes mythology — a fitting frame for a system that may occupy a qualitatively different tier of capability than what is currently publicly available.

    Technical Details

    While Anthropic has disclosed little about Mythos’s architecture, the framing of the leak offers some clues. The phrase “step change” is notable because Anthropic has historically been measured in its claims about capability improvements. The company’s Constitutional AI methodology and Responsible Scaling Policy (RSP) mean that any model flagged internally as a step change would likely trigger additional evaluation protocols before deployment — potentially including extended safety assessments, red-teaming exercises, and consultations with external researchers.

    Anthropic’s RSP defines AI Safety Levels (ASLs) that require progressively more stringent safeguards as models approach capability thresholds related to weapons development assistance, cyberoffensive potential, or autonomous self-replication. A model described internally as a step change in power would almost certainly be evaluated against ASL-3 and possibly ASL-4 criteria, the latter of which triggers a requirement that Anthropic demonstrate the model’s risks are adequately contained before commercial deployment.

    The cybersecurity concerns Anthropic flagged may relate to the model’s ability to generate novel attack techniques, assist in vulnerability discovery at scale, or operate in agentic settings with greater independence than prior Claude models. These are capability categories that the broader AI safety community has identified as particularly consequential as language models become more powerful.

    Industry Impact and Reactions

    The emergence of Mythos adds another dimension to an already turbulent period for Anthropic. The company is simultaneously navigating its lawsuit against the Trump administration over a Pentagon supply chain risk designation, an accelerating commercial subscription base, and a reported consideration of an IPO as early as October 2026. A breakthrough model — even one that remains internal — strengthens the company’s hand across all of these fronts, signaling continued technical competitiveness.

    AI researchers and industry observers noted that the leak itself is significant beyond the model’s existence. The fact that Anthropic felt compelled to confirm the disclosure while flagging new categories of cybersecurity risk suggests the company is actively managing the information environment around its most sensitive research, a posture that could become more common as AI labs push toward ever-higher capability tiers.

    Competitors will take note. OpenAI has been rapidly iterating its GPT-5 series, Google is pushing Gemini Ultra and custom AI chips, and Meta just launched its open-weight Llama 4 family. A Mythos-class model from Anthropic — if it achieves the step change described internally — would reset the competitive benchmark landscape in the second half of 2026.

    What Comes Next

    Anthropic has not announced a release date for Mythos, and industry analysts expect a lengthy evaluation period given the cybersecurity concerns the company has raised. Under Anthropic’s own RSP, any model triggering elevated risk assessments must pass a structured review before deployment. That process could take several months, meaning Mythos may not reach enterprise customers until late 2026 at the earliest — though limited research previews or staged rollouts to trusted partners remain possible.

    The company is also likely to face pressure from investors and the broader AI policy community to be transparent about the nature of the cybersecurity risks identified. As AI capability disclosures become an increasingly important part of the regulatory conversation in Washington and Brussels, Anthropic’s handling of the Mythos situation will be watched closely.

    Conclusion

    The accidental exposure of Anthropic’s Mythos model is a reminder that the frontier of AI capability is advancing faster than the public discourse typically reflects. With a model described internally as a step change now confirmed, and unprecedented cybersecurity concerns attached to its development, Anthropic faces the complex task of managing a breakthrough responsibly — even before it reaches users. How the company navigates the Mythos reveal may shape expectations for how advanced AI labs handle capability disclosures for years to come.

    Stay updated on the latest AI news at Evolve Digital.

  • OpenAI Releases GPT-5.4, Its Most Advanced Financial Reasoning Model Yet

    OpenAI Releases GPT-5.4, Its Most Advanced Financial Reasoning Model Yet

    OpenAI released GPT-5.4 on March 10, 2026, marking a significant step forward in the company push to make its models indispensable for high-stakes professional workflows. The latest model is designed specifically to excel at the kinds of complex financial analysis that typically require hours of expert work, and it arrives alongside a suite of new tools aimed squarely at enterprise finance teams.

    What Was Announced

    GPT-5.4, released in its Thinking variant, is now available across ChatGPT, Codex, and the OpenAI API. The model has been optimized with direct input from industry practitioners to improve performance on real-world finance tasks including financial modeling, scenario analysis, data extraction, and long-form research. OpenAI described it as the most capable model for financial reasoning the company has ever released.

    Alongside GPT-5.4, OpenAI announced ChatGPT for Excel in beta — a first-party Excel add-in that can build, update, and analyze financial models directly within workbooks. The integration adds financial data connections and uses GPT-5.4 Thinking to streamline workflows that analysts often spend days completing manually. The Excel add-in represents OpenAI first deep integration with Microsoft Office productivity software, extending the partnership between the two companies into everyday enterprise financial tools.

    A third announcement rounded out the release: Codex Security, an application security agent now available in research preview to ChatGPT Pro, Enterprise, Business, and Education users. Codex Security performs automated code vulnerability analysis, promising high-confidence findings, context-driven validation, and actionable remediation suggestions.

    Technical Details

    GPT-5.4 represents the latest in OpenAI incremental series of GPT-5 releases, each tuned for specific domains and use cases. The Thinking variant enables chain-of-thought reasoning, allowing the model to break down multi-step problems before producing a final answer — a technique that has proven particularly valuable for tasks like financial modeling, where accuracy and logical consistency are critical.

    The Excel integration works as a native add-in, embedding directly into the Microsoft Office environment rather than requiring users to switch between applications. This approach allows GPT-5.4 to access spreadsheet data in context, generating formulas, projections, and scenario analyses based on the actual content of open workbooks. Financial data integrations allow the model to pull in external data sources alongside local spreadsheet content.

    Codex Security, meanwhile, applies similar reasoning capabilities to the domain of software security, scanning codebases for vulnerabilities and generating detailed reports with specific remediation steps. The research preview targets organizations already using ChatGPT for development workflows who want to layer security analysis into their pipelines without adopting a separate tool.

    Industry Impact and Reactions

    The finance-first positioning of GPT-5.4 signals a strategic priority for OpenAI in enterprise revenue. Financial services has historically been one of the largest buyers of specialized AI tools, and embedding GPT-5.4 into workflows that analysts already rely on — particularly Excel — is a calculated move to make displacement of the model from those workflows difficult once adoption takes hold.

    The Excel integration in particular has attracted attention from enterprise technology analysts. Microsoft and OpenAI partnership has evolved steadily since OpenAI first took Microsoft investment, and direct integration with Microsoft 365 productivity tools like Excel represents a meaningful deepening of that relationship. Competitors including Google and Anthropic have each been building similar integrations with their own productivity suites.

    Codex Security arrives as enterprise demand for AI-assisted security tooling continues to climb. The research preview status keeps expectations measured, but the move into application security represents OpenAI expanding Codex beyond pure code generation into the governance and risk management side of software development.

    What Comes Next

    ChatGPT for Excel is currently in beta, with general availability timing not yet announced. OpenAI is expected to expand GPT-5.4 access across additional professional domains as the model moves out of initial release. Codex Security is in research preview and will likely evolve based on enterprise feedback before a broader rollout.

    The GPT-5 series has been releasing in rapid succession since the base model launched, and further refinements — potentially including GPT-5.5 — are expected in the coming months as OpenAI continues iterating on the frontier model line.

    Conclusion

    GPT-5.4 marks OpenAI ongoing effort to translate raw AI capability into tools that fit directly into professional workflows. By targeting financial reasoning and Excel integration together, OpenAI is betting that the path to enterprise stickiness runs through the spreadsheet — one of the most durable productivity tools in existence. Whether the strategy pays off will depend on how quickly finance teams adopt and depend on models they might not fully control.

    Stay updated on the latest AI news at Evolve Digital.