Jump to content
  • Sign Up
×
×
  • Create New...

ChatGPT

Diamond Member
  • Posts

    941
  • Joined

  • Last visited

  • Feedback

    0%

Everything posted by ChatGPT

  1. Nvidia has put nearly US$50 billion into the AI labs that buy its chips, and has lined up commitments for more than $500 billion Colette Kress, the company’s chief financial officer, told analysts on August 26 that demand from the labs Nvidia backs with its own balance sheet will contribute toward roughly a quarter of its business next year. That arrangement is what people mean by circular financing, and Nvidia used the phrase before any analyst did. Kress said on the earnings call that the company recognised the scale of the support it was providing and knew some would call it circular financing. She said Nvidia sees it differently. The loop is simple to describe. Nvidia invests in an AI lab. The lab uses the money, or the credit Nvidia’s involvement unlocks, to build a data centre. The data centre is filled with Nvidia chips. The purchase is recorded as Nvidia revenue. Nvidia’s share price and cash pile grow, and it invests again. What Nvidia has committed Kress gave the figures herself. She said Nvidia has signed partnerships with six investment firms, naming Apollo Global Management, BlackRock, Blackstone, Brookfield Asset Management, Goldman Sachs and KKR, to set up financing platforms that will raise more than $500 billion of outside capital for the labs to build with. She said Nvidia secured land, power and building capacity with SB Energy that will host only Nvidia equipment. The first phase supports 4.25 gigawatts and will be used by OpenAI. Kress put OpenAI’s existing and planned commitments at around 12 gigawatts of Nvidia compute through 2030. For a second lab she did not name, Nvidia will provide credit support covering nearly two gigawatts. Nvidia’s results statement describes those partnerships as subject to definitive agreements, which means the binding contracts have not been signed. The $500 billion is an intention rather than money in place. Nvidia is also lending its name to smaller cloud operators. Kress said the company promises to rent a portion of an operator’s capacity itself, which gives the operator’s lenders a guaranteed income stream to lend against. In exchange, Nvidia takes a share of what the operator earns above that floor. She said Nvidia gets paid twice under the arrangement, once on the equipment and again on the rental income. Why Nvidia rejects the circular financing label Kress gave three answers, and they deserve to be reported alongside the numbers. Outside lenders still assess every deal on its own merits, she said, and Nvidia is not making loans. The chips Nvidia ships go to customers that are investment grade or backed by someone who is. And if a customer does fail, the equipment can be moved to another buyer, which she offered as the reason Nvidia’s exposure is limited. Kress also explained why the labs need the help. They have more demand for computing than their finances can support, she said. They are young companies without the long contracts and credit ratings that lenders normally require before funding a data centre. What limits their growth is not customers or technology. It is access to computing. The obvious risk is what happens if one of them cannot pay. Nvidia would lose the ***** and the investment at the same time. Kress answers that the hardware finds another buyer. That claim only holds while demand exceeds supply, and Nvidia says it currently does. Vivek Arya of BofA Securities asked Jensen Huang how the company squares funding labs that are designing their own chips, pointing to OpenAI’s Jalapeño processor. Huang said Nvidia sells a platform that works in any cloud across the whole life of an AI system, while rival chips are built for one service. On the money, he said his only regret was not investing more and sooner. The agent assumption underneath it all Kress told Morgan Stanley’s Joseph Moore that an agent needs somewhere between 15 and 100 times the computing power of a person using the same system. Huang said he believes AI tipped over to being mostly agentic in the past month. Nvidia offered no data for that. On that basis, the company guided to $108 billion in revenue this quarter and said it preliminarily expects around 70% growth in the year to January 2028, a figure Kress said is limited by supply rather than demand. Kress separately warned that memory prices are climbing faster than Nvidia expected. She guided margins down to 74% this quarter, bottoming at 71% to 72% in the fourth, and said memory scarcity is being driven in large part by the AI buildout itself. Nvidia reports again on November 17. Kress did not say which lab is receiving the credit support covering nearly 2 gigawatts. (Photo by Nvidia) See also: NVIDIA Jetson Orin Nano 2 brings physical AI to drones and robots Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post A quarter of Nvidia’s business next year comes from labs it is financing appeared first on AI News. View the full article
  2. Autonomous trucking company Gatik has raised $200 million in Series D funding to expand its driverless freight operations across North America. The round was led by Qatar Investment Authority and Koch Disruptive Technologies, with participation from Millennium Management, ARK Invest, Intact Private Capital, and other investors. Gatik said it has more than $600 million in contracted revenue and has completed 85,000 fully driverless orders. The company reported a 99% on-time delivery rate across its operations. Its trucks currently move goods between distribution centres and stores across regional networks in Texas, Arizona, Arkansas, and Canada. Gatik operates dozens of fully driverless trucks and says it plans to expand the fleet to thousands over the coming years. A Gatik spokesperson told Reuters that the company is targeting more than 100 driverless trucks by the end of 2026. Gatik plans to use the new capital to expand commercial operations and its fleet while continuing to invest in technology, infrastructure, and its workforce. From fixed routes to dynamic freight networks Gatik’s autonomous vehicle system is designed for regional freight routes covering highways and surface streets. Gatik describes Gatik Driver as its proprietary AI system for autonomous freight. The company’s fully driverless trucks operate without a human driver or safety observer on board. When Gatik and Loblaw announced their initial Toronto fleet in 2020, five vehicles were scheduled to operate on five predetermined routes with fixed pickup and drop-off locations. By 2022, Gatik’s fully driverless Loblaw operation was transporting ambient, refrigerated, and frozen goods from a distribution facility to five nearby stores. Loblaw described the routes as fixed, repetitive, and predictable. PepsiCo said Gatik’s newer deployments can operate across highways and surface streets using dynamic route orchestration for networks spanning hundreds of pickup and drop-off locations. PepsiCo said route plans can be adjusted by adding or removing stops, responding to shifts in demand, and adapting to activity at distribution centres. The company said these changes can be made within its existing transportation operations without requiring major alterations to the network. TechCrunch reported that Gatik began with fixed trips of less than 10 miles but now operates dynamic routes containing dozens of pickup and drop-off points and covering distances of up to 400 miles. The company has developed its technology around middle-mile freight, where goods move between facilities such as warehouses, distribution centres, and retail locations rather than directly to consumers. In June 2026, Gatik signed a multi-year agreement with PepsiCo to deploy autonomous freight vehicles within the company’s North American supply chain. The trucks are operating across Texas, Arizona, and Arkansas. The agreement focuses on regional transportation networks where products move between sites on high-frequency schedules. TechCrunch reported that the PepsiCo operation includes 41 fully driverless box trucks moving Frito-Lay products between distribution centres and stores in Dallas, Phoenix, and northwest Arkansas. PepsiCo said its first deployment with Gatik began in 2022, several years before the companies expanded their commercial arrangement in 2026. Gatik has also expanded its operations with ********* retailer Loblaw. In September 2025, the companies signed a five-year agreement covering an initial deployment of 50 autonomous trucks across Loblaw’s distribution network in the Greater Toronto Area. Under that plan, 20 trucks were scheduled for deployment by the end of 2025, followed by another 30 by the end of 2026. The vehicles are intended to serve more than 300 Loblaw stores and transition from operations with safety drivers to fully driverless freight service. Gatik and Loblaw began autonomous deliveries in 2020 before removing the safety driver from commercial routes in 2022. Before the fully driverless deployment, the companies said they had completed more than 150,000 autonomous deliveries with a safety driver on board. Scaling AI and autonomous truck production Gatik also uses simulation and synthetic data to develop and validate its autonomous driving software. In July 2025, the company introduced Arena, an internally developed simulation platform designed to reproduce driving environments without relying exclusively on physical road testing. Arena generates structured synthetic data that can be used to test autonomous driving behaviour across different conditions. The platform is designed to reproduce routine situations as well as rare or high-risk scenarios that are harder to encounter repeatedly during real-world testing. Arena incorporates Nvidia Cosmos world foundation models to generate synthetic driving environments used for training and validation. The platform can vary driving conditions and reproduce scenarios that occur infrequently during commercial operations, allowing Gatik to test the same situations repeatedly in simulation. Gatik also uses Nvidia hardware for onboard AI processing. Nvidia said in March 2025 that Gatik was integrating DRIVE AGX into its Class 6 and 7 autonomous trucks for real-time processing of the data required by its autonomous driving system. Nvidia describes DRIVE AGX as the onboard computing platform used to process real-time sensor data and run AI workloads for autonomous driving. Isuzu Motors invested $30 million in Gatik in 2024 as part of a partnership to develop Level 4 autonomous commercial vehicles in North America. The companies are jointly developing a redundant chassis designed for autonomous driving. Isuzu said the redundancy is intended to provide a means of recovery if the autonomous driving software experiences a defect. Isuzu and Gatik are targeting mass production of the jointly developed chassis from 2027. Isuzu is also targeting 2027 for the start of its Level 4 autonomous commercial vehicle business. Isuzu’s 2025 integrated report lists Gatik as its partner for autonomous logistics using medium-duty trucks in the North American middle-mile market. The collaboration covers both autonomous vehicle development and mass production of chassis designed for autonomous systems. Gatik’s trucks are classified as Level 4 autonomous vehicles. At Level 4, the automated driving system can perform the driving task without human intervention when the vehicle is operating within the conditions for which the system has been designed and validated. Those operating limits are commonly defined through an Operational Design Domain, or ODD. Isuzu defines an ODD as the road, geographic, environmental, and other conditions that must be satisfied before an autonomous driving system operates. Gatik’s current third-generation trucks operate on both highways and surface streets. CEO Gautam Narang told TechCrunch that the vehicles can operate around the clock and handle conditions including light rain and snow. (Photo by Bernd Dittrich) See also: Amazon’s Prime Air autonomous drones to reach 500 US cities Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post Gatik raises $200M to scale AI-powered autonomous freight appeared first on AI News. View the full article
  3. NVIDIA has unveiled the Jetson Orin Nano 2, an edge robotics computer aimed at bringing physical AI to drones, robots, and vision systems. The company is positioning the new board as an entry-level option for developers who want generative AI models running directly on a machine instead of inside a data centre. NVIDIA’s argument for the launch rests on a change in how well small and medium AI models now perform. The company says models of this size have reached the accuracy that only the largest frontier models achieved a year earlier. Deepu Talla, VP of Robotics and Edge AI at NVIDIA, said: “The Jetson Orin Nano 2 computer puts that breakthrough within reach of millions of developers, delivering the performance and energy efficiency needed for real-time reasoning at the edge.” Achieving that level of accuracy from small and medium AI models lets compact edge hardware interpret language and images and act on that information in real-time. Robots, delivery drones, inspection drones, and vision AI systems all depend on hardware that can run those workloads without drawing much power. NVIDIA Jetson Orin Nano 2 compute specs and power draw Jetson Orin Nano 2 carries 78 trillion operations per second of AI compute, 8GB of memory, and an eight-core Arm CPU. NVIDIA built the board to deliver a jump in AI and video-processing performance while keeping cost and power draw low. The new board reaches twice the inference performance of the existing Jetson Orin Nano Super. NVIDIA attributes that gain to improved Tensor Cores and higher memory bandwidth, packed inside the same compact form factor as its predecessor. Running in 15-watt mode, Jetson Orin Nano 2 uses 40 percent less power than the Orin Nano Super while matching its performance level. Jetson Orin Nano 2 runs on NVIDIA’s open software stack alongside Jetson agent skills and the wider Jetson AI ecosystem. NVIDIA says the board is built to run large language models and vision language models optimised for memory-efficient inference at the edge. The company names its own Cosmos and Nemotron models as examples, alongside Gemma 4 and Qwen 3, as models developers can deploy on the hardware. Early partners test physical AI applications Cognex, Doosan Bobcat, and Matic sit among the first companies NVIDIA names as adopting and exploring Jetson Orin Nano 2. NVIDIA says more than three million developers already build on its robotics stack. The company expects partners – including the aforementioned partners – to bring edge AI into home robots, vision AI systems, delivery and inspection drones, carrier boards, hardware systems, and reference designs. Wing, the drone delivery subsidiary of Alphabet, already runs Jetson Orin Nano Super and NVIDIA’s software stack across its delivery drone fleet. The company plans to evaluate Jetson Orin Nano 2 to push further into real-time AI perception and reasoning, with the aim of making deliveries from local businesses to residential yards faster and safer. Dinuka Abeywardena, Head of Perception at Wing, commented: “Drone delivery depends on AI that can enable fast, reliable understanding of the real world. Wing is exploring Jetson Orin Nano 2 to give us a path to more responsive, energy-efficient drones that can help make deliveries quicker and more dependable for customers.” Wing’s existing drone fleet runs on the Jetson Orin Nano Super, a separate product from the new board. There’s currently no public timeline for Wing to move from evaluation to production use of Jetson Orin Nano 2 on delivery flights. Matic Robots, a consumer robotics company, is adopting Jetson Orin Nano 2 for its home cleaning robots. NVIDIA says the board will let Matic add conversational AI, gesture detection, precision mapping and semantic understanding of the home, alongside autonomous cleaning behaviour. Navneet Dalal, Cofounder and Chief Executive of Matic Robots, explained: “Home robots need to understand people, map spaces precisely, understand the layout of objects and spaces, and clean autonomously in dynamic and constantly changing environments. “With Jetson Orin Nano 2, Matic can run state-of-the-art AI models at the edge in a compact home robotics platform built for real-time perception, interaction, and navigation.” Carrier boards and reference designs NVIDIA also named a large roster of Jetson hardware partners building around the new board. AAEON, ADLINK, Advantech, and Aetina sit among the manufacturers building carrier boards and hardware systems for it. Antmicro, Aptiv, Auvidea and AVerMedia join a further group working on customised AI software and reference designs, alongside Chuanglebo, Connect Tech, ForeCR, and JWIPC. Neurealm, Plink, Realtimes, RidgeRun, RS, Seeed Studio, Tauro Tech, Twowin, TZTEK, and YUAN complete the list of partners NVIDIA says are working to help customers reach the market faster. “Frontier intelligence has reached the edge. Frontier models that used to run inside data centers last year can now run in real time on entry-level Jetson systems,” said Talla. “With Jetson Orin Nano 2, robots and edge systems can run leading language and vision models locally and in real-time, opening up applications that weren’t possible before.” Learn more about physical AI during the Physical AI Expo held in Amsterdam, London, and North America. See also: XPENG IRON humanoid robot draws record physical AI funding Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post NVIDIA Jetson Orin Nano 2 brings physical AI to drones and robots appeared first on AI News. View the full article
  4. MIT engineers have built an AI tool that forecasts extreme weather without training on historical disaster data. Kai Chang, a mechanical engineering graduate student, and Professor Themis Sapsis developed the tool. It produces maps of events that have not appeared in a region’s historical record but remain statistically-possible. Each map also carries estimates of the event’s likely duration and intensity, alongside a separate estimate of the area it might affect. Forecasting extreme weather events without historical precedent Sapsis holds the William I. Koch Professorship in Mechanical and Ocean Engineering at MIT. Both researchers are affiliated with the MIT Center for Computational Science and Engineering, and Sapsis also holds an appointment with the MIT Institute for Data, Systems, and Society. The pair describe the method, named Extreme Event Aware or η-learning, in a paper published in Nature Communications on 20 August. Existing risk models work differently. Insurers, city planners, and grid operators typically want to know what a once-in-a-century storm might look like for a specific location. Current simulations usually depend on datasets that already contain extreme events, learning the conditions that produced them before projecting similar patterns forward. Chang argues this current approach creates a limit on what such models can show. “These methods assume there are very disastrous events that we have seen in the dataset, and they build a method to either estimate the risk of those events, or they try to predict exactly the events that have happened,” he says. Sapsis frames the same limitation through Hurricane Katrina. “An event like Hurricane Katrina is something that happens every 30 to 40 years,” he adds. “What will be the Katrina that happens every 100 years? How bad will it be? That’s exactly what we’re trying to quantify, to help planners prepare for plausible extreme scenarios.” Combining point statistics with spatial detail The algorithm works from two types of data. Point statistics capture how often a given intensity level, such as the maximum rainfall recorded across a map, occurs within a dataset. Spatial maps show how an event’s impact varies across a region. Learning the statistical relationship between the two lets the algorithm build spatial patterns for events beyond anything in its training data, without needing prior examples of those exact extremes. The researchers tested the approach on precipitation across the continental US. They started with 25 years of hourly rainfall data, pooled into daily maps, and computed point statistics describing how often the maximum rainfall on a map reached a given level across that full record. The training window for the spatial model was narrow. They trained that part of the algorithm using paired low-resolution and high-resolution maps drawn from only the first six months of the 25-year record, a ******* that contained few or no examples of the heaviest rainfall levels. The algorithm learned how patterns in the low-resolution maps corresponded to detail in the high-resolution versions, then applied the point statistics from the full record to constrain how extreme the generated patterns could become. Testing infrastructure against worst-case maps The highest rainfall ever recorded in New York City measures 200 millimetres. The method can generate plausible maps of a storm that produces 300 millimetres instead, a level with no match in the observational record. A user can prompt the trained algorithm to show what a once-in-a-century storm might look like for a named city. The output takes the form of maps showing statistically-plausible storms at that frequency. Each map carries its own size and area of coverage, and rainfall intensity varies across the set as well. According to Chang, the algorithm can generate large volumes of these scenarios at once. The generated maps could help a city test its seawall against a storm surge beyond anything recorded. The same maps could show whether the power grid would hold during a longer heatwave, or whether firefighting resources could contain a wildfire larger than any on file. Limits of the demonstration so far Applying the method to a new hazard requires relevant point statistics and spatial data for that specific hazard, according to Chang and Sapsis. The pair point to possible extensions once that data is available, such as visualising severe floods and wildfires with no equivalent in the historical record. Sapsis notes that global infrastructure has been optimised for efficiency, leaving little slack in the systems it supports. “A single extreme event propagates through supply chains, energy markets, and food systems in weeks,” he explains. “Being able to put a probability on an event that hasn’t happened yet is now a question of national and economic resilience.” See also: Samsung health AI models analyse wearable biosignal data Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post MIT AI forecasts extreme weather without historical data appeared first on AI News. View the full article
  5. XPENG’s physical AI unit has secured over $900 million at a $6.3 billion valuation to scale its IRON humanoid robot platform. The ******** electric vehicle maker announced the funding round for its robotics business through a set of share purchase agreements with multiple investors. XPENG says the deal represents the largest single-round private capital raise in China’s physical AI industry to date, based on both the amount raised and the resulting valuation. He Xiaopeng, Chairman and CEO of XPENG, said: “Over the past 12 years, XPENG has remained committed to full-stack in-house R&D, building a solid technological foundation for the physical AI era—across our physical world foundation model, Turing AI chips, and AI infrastructure. “This has enabled us to pioneer a new phase of mass production and commercial deployment for advanced humanoid robots.” However, Xpeng’s shares are down 7.22 percent at the time of writing, and the 12-month decline stands at 51.24 percent. XPENG’s core vehicle business faces intensifying competition inside China from domestic manufacturers, and from Tesla in both overseas and domestic markets. Investors are pricing the robotics division at $6.3 billion while the parent company’s shares keep falling. IDG Capital leads the round, Tencent and Alibaba join as backers IDG Capital led the financing round, with Gaorong Ventures also participating as an investor. Tencent and Alibaba joined as strategic investors. XPENG will keep controlling ownership of the robotics business once the round closes, and the unit will remain consolidated into the group’s financial statements. XPENG said the capital will fund long-term investment in what it terms full-stack physical AI development. The company also plans to use the raise to strengthen incentive arrangements for senior executives and other staff working on robotics. According to XPENG, the money will support software and hardware R&D for the unit, plus training and iteration of its physical AI models. Further funds are earmarked for high-quality data generation and for building end-to-end mass production facilities. XPENG also intends to put some of the capital toward commercial expansion outside China. Inside IRON, XPENG’s humanoid robot platform IRON is at the centre of XPENG’s physical AI strategy. The humanoid robot uses a fully-enclosed flexible lattice structure that XPENG designed in-house to balance appearance with safety. IRON has 76 degrees of freedom across its body and 21 in each hand, which XPENG presents as evidence of its robot’s high dexterity and mobility. XPENG built the robot’s hardware platform itself, including the chips and controllers that drive the robot’s core movement systems. Separate motion modules and dexterous hand mechanisms handle finer manipulation tasks. The company is applying quality standards and production processes developed for its electric vehicles to robot manufacturing, aiming for automotive-scale output and delivery volumes. On compute, XPENG puts IRON’s combined output at up to 2,250 TOPS of effective computing power, delivered across three in-house-designed Turing AI chips. That on-board processing lets XPENG run its physical AI foundation model directly on the robot. IRON can then carry out complex tasks with low inference latency and with data processed locally on the device. XPENG argues that IRON’s human-like hardware gives it an advantage in collecting behavioural data from everyday human activity, and in adapting to environments and tools built for people. As IRON reaches mass production, the company expects what it calls a data-model-application flywheel to take hold, accelerating the robot’s ability to learn and take on new tasks. IRON is expected to enter mass production by the end of 2026. Initial deployment will happen inside the company’s own stores and campuses before any wider rollout. XPENG plans to begin deliveries to customers in China and overseas markets during 2027. Investors point to XPENG’s physical AI manufacturing scale IDG Capital said the physical AI industry is moving “from technical breakthroughs to scalable manufacturing and commercial deployment,” a shift it argues plays to XPENG’s strengths. The firm pointed to XPENG’s combination of edge AI processors, physical AI foundation models, and complete robotic systems, and to the technology links between the robotics unit and XPENG’s electric vehicle and autonomous driving businesses. Gaorong Ventures added that humanoid robots are “moving beyond demonstrations of mobility and dexterity toward reliable mass production and tangible value creation in real-world settings.” The firm highlights XPENG’s automotive-grade safety and quality standards as evidence the robot is built for commercial deployment and credited XPENG’s decade-plus of supply chain and manufacturing experience in the EV sector. He Xiaopeng commented: “I believe the strong capital backing from leading global and strategic investors provides the resources needed to accelerate the growth of our robotics business, while strengthening our ability to attract more world-class physical AI talent. “IRON brings together a highly human-like design, advanced AI intelligence, built to the highest standards of safety and quality. Our ambition is for IRON to become a trusted partner for people and a meaningful part of everyday work and life.” Learn more about physical AI during the Physical AI Expo held in Amsterdam, London, and North America. See also: Amazon’s Prime Air autonomous drones to reach 500 US cities Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post XPENG IRON humanoid robot draws record physical AI funding appeared first on AI News. View the full article
  6. In August 2025, TypeScript became the most used language on GitHub. This was the largest shift in GitHub’s language rankings in the last ten years and it occurred during the ******* of most accelerated adoption of coding AI agents. Coding AI agents had previously been predicted to lower the importance of language selection. It was assumed that organizations would become stack agnostic and select tech stacks based solely on the needs of the business problem, leaving behind the considerations of the available developer hiring pool. Instead, a mere two years after the widespread adoption of AI coding tools, the market appears more restricted, with a rapid narrowing of available coding languages, and most of the focus is on a single family of languages. Analysis of the change GitHub’s October 2025 Octoverse report counted TypeScript contributors. Its 2.64 million monthly contributors marked a 66% increase year-over-year. During 2025, TypeScript saw over a million developers write their first TypeScript code in GitHub. This growth is in addition to an already highly dominant position. In 2025 Stack Overflow sent a developer survey that collected more than 49,000 responses. Of these responses, 66% self-reported using JavaScript. For nearly every year since 2011 JavaScript has dominated this position. In conclusion the JavaScript family is both the most used and the fastest growing on GitHub. It is worth briefly addressing an issue with GitHub’s counting method; counting activity on GitHub is counting activity on their own site which presents a conflict of interest as they may want it to look good. Also, fashion trends impact the kind of information that is included in public repositories. However, the indicators still align with survey data, which is reasonable considering the scope of this particular trend. Models write best in the code they have seen most How it works is quite simple. Models learn from the code that is published, and most of the code that is published is written in JavaScript and TypeScript. In fact, much of the code is centered around React. This creates a substantial gap in the output that developers can see in their agents within a day of changing their stack. Request a coding agent to generate a typed React component and the output will usually compile, conform to the standards of the codebase, and require minimal edits. In contrast, when you ask the same agent to generate code for a Svelte, Solid, or a less popular backend framework, the output tends to be much thinner. Also, you will see more invented APIs, and the scaffolding will need more corrections before it is executable. As a result, this is now changing the way that teams are selecting their stacks. It is no longer simply a matter of deciding which framework is the most efficient or easiest to work with, but rather which framework is most compatible with the team’s tools because the productivity gap in usable agent output during an extended build cycle is compounded. When that team produces output, it gets published, scraped, and included in the subsequent training runs, widening the productivity gap. None of this points to a technical verdict. Solid and Svelte are good frameworks, and several more modern frameworks outperform React on raw speed. The market rewarded the choice the models already knew. The models are built in Python, but the products are shipped in JavaScript A valid counter to the above points is that AI development happens in Python. Model training, evaluation, and most research tooling run in Python, and that hasn’t changed. However, very little of what the customer interacts with is written in Python. In fact, the front end of an AI product is essentially a window that streams tokens. It also requires buttons to execute tools, an approval step for anything that could lead to a negative outcome, and an explanation of what the system did and why. This is all done in JavaScript and TypeScript, regardless of whether the model has been sourced from OpenAI, Anthropic, or an open-weight model that the company hosts on their own hardware. By the end of 2025, GitHub had reported over 1.1 million public repositories using an LLM SDK, a 178% increase from the previous year. This increase was primarily due to application development, rather than model development. Every enterprise pilot that makes it past the demo stage requires someone to build the user-facing component, and the industry-standard tools for that work are JavaScript frameworks. Type systems became the guardrail for generated code The opinion from GitHub is that the shift has meant developers are moving towards typed languages because type systems make agent-assisted development safer. Generated code has a specific type of failure. It reads and is structured well. It even runs perfectly fine in dynamic languages, only to ****** due to shape mismatches three calls down. Type checkers will identify a large portion of this before the code even gets run. This theory has proven correct in practice. In 2026, the use of TypeScript by professional developers reached 78 percent, an increase from 69 percent two years prior. Approximately 40 percent of developers write exclusively in TypeScript, and only 6 percent of developers write exclusively in plain JavaScript. The types of errors that a compiler will find tend to be the types of errors that a human reviewer will overlook when faced with 400 lines of reasonable code. ● A function is called with an object that is missing one of the required fields. ● Code contains an assumption that a value exists, and this results in a null or undefined value being passed. ● The shape of an API response has changed, and the generated handler still uses the old one. The bottleneck has changed from writing code to verifying code The field report from OpenAI regarding the use of coding agents in scientific computing from July 2026 stated the limitation most plainly: verification has become the limiting factor, as opposed to code generation. While it is a vendor examining their own product and should be taken with a grain of salt, the result coincides with what many engineering teams, particularly those outside of research, have been noting for the past year. When a capable front end is finished in one afternoon instead of three weeks, the slowest part of the process becomes determining if everything that appeared on the screen is correct, secure, and maintainable. This changes what is expected from a JavaScript developer. Speed in typing code has never been the real value of the job, but it was used as a rough measure of competence during the hiring process. Generated code has taken away the measure and left the judgment. It is expected that a React effect will fire twice during development, and a developer who is unaware of this may spend an entire day investigating what they think is a bug due to a duplicate API call. A generated query will seem fine in the development environment with sample data, but it can end up scanning an entire table when used in the production environment. An auth check can be placed anywhere in a component and be ineffective, which may give the illusion of security, but that illusion disappears when someone actually tests it. There is a misalignment that teams tend to overlook. Generation capacity is nearly infinite and increases with each additional agent or a new subscription. In contrast, review capacity is limited by the number of engineers who are sufficiently versed in the system to identify a plausible error. That number is not likely to increase at the same rate. Adding more code generation capacity to a team that is already at their review capacity limit does not increase the rate of delivery. It simply moves the bottleneck from writing to reviewing. A team could double their code generation capacity in a week, but that won’t change the number of people available to review. This is why the constraint has shifted and why additional tooling does not provide the solution. The same has not been true for hiring practices. Most screening still assesses whether a candidate can arrive at a working solution, which is the part the tools already assist with. Some companies have begun evaluating the opposite skill. They present candidates with blocks of AI-generated code which contain a fault and observe how long the fault remains unaddressed. Staffing firms have also moved in that direction. For instance, Full Scale now describes the JavaScript engineers it places as having fluency in AI tools and product sense, rather than lines of code written. A company that sets out to hire a dedicated JavaScript developer is now acquiring as much review capability as building capacity. That seems a minor shift until an AI-generated login flow moves into production without human intervention. The concentration carries a cost A market that values what the models already know makes it difficult to introduce anything new. A new framework published this year has no pre-existing corpus of training data, making it difficult for agents to work with it, causing teams to avoid it, leading to a scenario where no corpus is generated. The typical time-to-funding is insufficient for merit-based solutions to break this cycle. Frameworks that achieved their milestones prior to 2023 now have an advantage that has nothing to do with quality of design. The risk is narrower for a single company. A business whose product, tools, and hiring pipeline all revolve around a single language family has made the same bet three times. That is comfortable while the language family maintains its dominance, and costly should it falter. The first prediction was half correct. AI has indeed eliminated much of the cost of writing code in a language that no one on the team understood. The cost of comprehension and ownership still persists, and for most teams, this is the primary cost that determines the technology stack. The post How AI coding tools are contributing to the popularity of JavaScript appeared first on AI News. View the full article
  7. Amazon plans to expand its Prime Air drone delivery service to nearly 500 cities and towns across the US by the end of 2026. That build-out amounts to six times the number of locations Prime Air serves today, extending the option to communities with tens of millions of customers, according to Amazon. Reaching that many locations without adding pilots to each flight depends on the drones’ own decision-making systems rather than a large ground staff monitoring individual flights. Prime Air’s fleet runs on what Amazon calls “highly autonomous” flight software, engineered to keep functioning safely and predictably when something unexpected happens mid-flight. A Detect-and-Avoid system sits at the centre of that setup, continuously scanning the airspace and surroundings around each drone much like a pilot checking for other aircraft. That scanning lets the drone spot obstacles on its own and make real-time flight decisions without a remote operator stepping in. Onboard cameras and sensors handle navigation, obstacle detection, and the delivery drop itself. Amazon has stated the cameras do not track individuals or record their movements, and the footage feeds only the drone’s own navigation processing rather than a monitored video feed sent back to base. No person watches a live camera stream from any of the aircraft, according to the company. FAA Part 135 certification supports the expansion Prime Air operates under Federal Aviation Administration Part 135 certification, the licence category used for commercial air carriers. Amazon points to this as the highest tier of FAA oversight available to a drone delivery operation, above a lighter drone-specific exemption. Holding that certification is what lets the fleet add new metro areas without needing a separate waiver for each one. The safety systems extend to landing behaviour under adverse conditions. Amazon says its advanced safety systems are built to bring a drone to a safe landing in the event of severe weather or other unexpected events, with the stated goal of protecting people, pets, and property on the ground. The drones are also engineered for everyday flying conditions rather than fair-weather use alone, including light rain and a range of temperatures, and they run fully-electric with zero exhaust emissions. Noise output factors into how Amazon positions the technology for residential areas. During drop-off, Amazon says sound levels sit below an idling delivery truck parked at the kerb and last around 30 seconds. At cruising altitude, the company compares the sound to a window fan on its low setting, typically inaudible indoors. Tiered logistics network built around speed Prime Air carries items weighing five pounds (~2.27kg) or less that fit in a large shoebox, a limit Amazon says covers more than 60 percent of the items customers most frequently buy on the platform. The eligible catalogue spans millions of items at Amazon’s standard pricing, including groceries, cosmetics, medications, and electronics alongside harder-to-find niche products. Specific examples Amazon lists include iPhones, Samsung Galaxy handsets, Apple AirTags and AirPods, Ring doorbells, and an Alpha Grillers instant-read food thermometer. Orders can land in as fast as 30 minutes, though Amazon puts the typical wait closer to 60 minutes after checkout. The fee structure ties to Prime status and order size: Prime membership gets drone delivery free on orders of $50 or more, a $2.99 charge applies to smaller Prime orders, and customers without a membership pay $4.99 regardless of basket size. First-time drone customers set a delivery point when they place that initial order, then can reuse it or pick a new spot for later purchases. Before releasing a package, the drone scans the delivery point for people, pets, or vehicles in the way. However, it doesn’t yet always get it right: Lindsey Austen, who lives in Richmond, Texas, had gotten an alert that her Amazon order would be delivered by drone to her home for the first time. Footage shows the drone hovering above her pool, then releasing her package into the water pic.twitter.com/yTtpUqv5XP — Sky News (@SkyNews) August 20, 2026 Prime Air occupies one tier in a broader set of Amazon delivery options rather than replacing them. Amazon Now offers 30-minute delivery in dozens of US cities on thousands of grocery, household, and locally-relevant items, dispatched from smaller sites concentrated in more populated areas. One-hour and three-hour delivery covers a selection of more than 90,000 items through Amazon’s Same-Day Delivery network. Same-Day Delivery itself reaches millions of items across more than 10,000 cities and towns, including a growing number of smaller towns and rural areas. Amazon views Prime Air as complementary to that logistics structure: drone delivery from Amazon’s own sites, positioned primarily for suburban areas, sitting alongside denser urban same-day options and the wider Same-Day network’s rural reach. Current footprint and what comes next 11 Prime Air sites now cover 10 metro areas across seven states. Arizona’s site sits in Tolleson, near Phoenix, and Florida’s sits in Ruskin, near Tampa. Kansas City and Baton Rouge each host a single site outright, and Michigan runs two – Hazel Park and Pontiac – both serving the Detroit area. Omaha’s operation is based in Papillion, Nebraska, and Texas alone accounts for four locations: Richmond near Houston, Richardson near Dallas, plus San Antonio and Waco directly. Each site covers roughly 175 square miles of surrounding territory. Amazon reports that Prime Air’s drone delivery sites post the highest average delivery volumes of any US drone operation, with thousands of deliveries made daily, and says it has already delivered hundreds of thousands of packages by drone this year. David Carbon, VP of Amazon Prime Air, said: “Customers already turn to Amazon for fast Same- and Next-Day Delivery, and Prime Air provides them an even speedier option when they need it, with deliveries in as fast as 30 minutes. “We’ve already delivered hundreds of thousands of packages to customers by drone this year, and by the end of 2026 we plan to reach customers in nearly 500 cities and towns. We’re pairing the convenience of ultrafast drone delivery with the low prices and selection customers love about Amazon.” Amazon plans to launch Prime Air service soon in the Chicago, Illinois; Syracuse, New York; Cleveland, Ohio; Atlanta, Georgia; and Boise, Idaho metro areas, with further communities scheduled to follow later this year. Coverage by ZIP code remains subject to change, and Amazon notes that not every address inside a listed ZIP code qualifies for drone service. The service has also moved beyond the US. Amazon’s Prime Air site in Darlington, *** is the first of its kind outside the US, giving the autonomous flight and Detect-and-Avoid systems their first operational test under a different national aviation regulator. Learn more about physical AI during the Physical AI Expo held in Amsterdam, London, and North America. See also: Alvys launches AI agents for freight TMS workflows Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post Amazon’s Prime Air autonomous drones to reach 500 US cities appeared first on AI News. View the full article
  8. Stripe has agreed to acquire OpenRouter, an AI model-routing platform that gives developers access to hundreds of models through a single interface. The deal adds model selection and routing to Stripe’s existing work around AI usage and token-based billing. OpenRouter supports more than 400 models from over 80 providers, according to Stripe. Rather than requiring separate integrations with each model provider, developers can use OpenRouter to send requests through one API. Routing beyond model choice The platform evaluates requests using factors including task complexity, price, speed, and reliability. It can then direct each request to a model suited to those requirements. OpenRouter also handles a second layer of routing between providers serving the same model. Its documentation says customers can prioritise endpoints based on price, throughput, or latency, while setting requirements such as maximum prices or minimum performance levels. The platform measures latency and throughput for individual model-provider combinations using rolling performance data. This allows a request to be routed to an endpoint that meets specified cost or performance criteria rather than relying on a fixed provider. That separates two routing decisions: which model handles a request and which provider endpoint serves it. Provider choice can also affect inference costs even when the underlying model remains the same. In a June 2026 example, OpenRouter listed Llama 3.3 70B input pricing at $0.10 per million tokens through DeepInfra and $1.04 through Together, while output pricing ranged from $0.32 to $1.04 per million tokens across the providers shown. Routing can provide failover when an endpoint becomes unavailable. OpenRouter says its system can move requests to alternative providers or models when it encounters problems including provider outages, rate limits, context-length errors, or moderation refusals. Data-handling requirements can also form part of provider selection. OpenRouter lets users restrict requests to Zero Data Retention endpoints and prevent routing to providers that collect data or train on prompts, while enterprise customers can request in-region processing in the US or EU. Routing criteria can therefore include model capability, provider availability, processing location, latency, throughput, and cost. Multi-model infrastructure expands Multi-model environments are already common among surveyed organisations. F5’s 2026 State of Application Strategy report, based on responses from more than 1,100 IT decision-makers, found that 52% of organisations were chaining or orchestrating multiple AI models, with respondents using an average of seven models. Menlo Ventures, an OpenRouter investor, reported a different measure of provider behaviour in its 2025 mid-year survey. It found that 66% of builders upgraded models while staying with their existing provider, while 11% switched vendors. OpenRouter is also one of several infrastructure providers adding model routing. Snowflake announced dynamic model routing for Cortex AI Gateway on August 18, with the feature expected to enter private preview. Snowflake said the system will assign requests according to factors including quality, speed, customer preferences, and cost. Cloudflare offers Dynamic Routing in beta through AI Gateway, with rules covering model selection, quotas, and fallbacks. AWS provides Intelligent Prompt Routing through Bedrock, while Microsoft Foundry offers routing profiles that balance model quality and price. AWS and Snowflake both describe systems that can direct less demanding workloads to smaller or lower-cost models while reserving other models for tasks requiring higher response quality or more complex reasoning. Routing also adds operational requirements. Microsoft’s Azure Architecture Center notes that dynamic model selection can complicate cost forecasting, debugging, and performance analysis when different requests are handled by different models. Stripe and OpenRouter were already working together before the acquisition. In January 2026, Stripe said developers using OpenRouter could route model requests through the platform while Stripe tracked usage, applied pricing, and handled billing. The arrangement paired OpenRouter’s routing layer with Stripe’s usage measurement and billing systems before the acquisition agreement. Token usage meets billing Stripe has also been developing token-based billing tools for AI applications. Its LLM token-billing service, which Stripe currently lists as being in private preview, can meter consumption according to model and token type, including input, output, and cached tokens where supported. Stripe’s documentation says businesses can use the system for per-token pricing, prepaid credits, fixed fees with included usage, or combinations of those approaches. The company can also update supported model prices when providers change their underlying pricing. OpenRouter already produces much of the usage data involved in those billing calculations. Its API reports prompt, completion, reasoning, and cached token counts with individual responses, along with the cost of the request. It also records the underlying inference cost charged by the provider separately from the amount charged to an OpenRouter account. Token counts are calculated using each model’s native tokeniser rather than applying a single counting method across all models. Stripe CEO Patrick Collison has linked the acquisition to the role of tokens and computing resources in AI applications. Collison said tokens are a central unit for companies building with AI and tied their economic use to how companies manage available computing resources. Enterprise token consumption is already reaching large volumes. Deloitte surveyed 515 US-based business and technology decision-makers in late 2025, all from organisations generating at least $500 million in annual revenue. The survey found that 37% of respondents were consuming between one billion and 10 billion AI tokens per month, while another 30% were consuming more than 10 billion. By 2028, 61% expect monthly consumption to exceed 10 billion tokens. Deloitte said much of the increase would come from workloads exceeding 100 billion tokens per month, with token use in that range expected to triple from 2026 to 2028. The firm cautioned that higher token consumption does not necessarily indicate more effective AI adoption. Deloitte identified oversized prompts, weak context management, and limited reuse as factors that can increase token consumption. The cost of processing those tokens can also differ by model and, in some cases, by the provider serving the same model. OpenRouter says it processes more than 10 trillion tokens per day across a community of more than 10 million developers and companies. Earlier figures published by OpenRouter provide some indication of how its traffic had changed before the acquisition. In May, the company said its weekly volume had increased from five trillion to 25 trillion tokens over the previous six months. At the time, OpenRouter reported serving more than eight million developers across more than 400 models. OpenRouter was founded in 2023 and has raised funding from investors including Menlo Ventures and Andreessen Horowitz. Its $113 million Series B round in May was led by CapitalG, Alphabet’s independent growth fund, with participation from investors including NVentures, ServiceNow Ventures, MongoDB Ventures, Snowflake Ventures, and Databricks Ventures. Stripe and OpenRouter did not disclose the financial terms of the acquisition. Reuters reported that the transaction is worth slightly more than $8 billion, citing a person familiar with the matter who requested anonymity because the information was confidential. (Photo by appshunter.io) See also: OpenAI president urges enterprises to hasten AI security defences Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post Stripe agrees to buy OpenRouter as AI model routing expands appeared first on AI News. View the full article
  9. The United Arab Emirates (UAE) has been early in adopting artificial intelligence for 9 years. It published a national AI strategy in October 2017 and, days later, created a ministerial post to run it, making Omar Sultan Al Olama the world’s first minister of state for artificial intelligence at 27. In the years since, it has built the digital identity, sovereign cloud and data-sharing layers most governments are still procuring. Agentic AI in government is where that head start either pays off or does not. The test began this month. More than 100 federal officials attended a workshop to launch the strategic track of the UAE’s agentic AI project, the national programme to convert half of federal government operations, services, and tasks to agentic AI models within two years. Organised by the National Committee for the Agentic AI Project, the session set implementation priorities and timelines, and began building the frameworks that will classify which government tasks can be handed to an agent at all. That classification exercise is the whole ballgame. Compute and plumbing the UAE has, including an AI-powered proactive performance system that tracks more than 150 million data points a month. What no government anywhere has yet published is a defensible method for drawing the line between a task an autonomous system may complete and one it may only recommend. Huda Al Hashimi, Deputy Minister of Cabinet Affairs for Strategic Affairs, described the workshop as the operational beginning of deploying AI models across government activity. The programme runs across seven pillars: strategy and projects, foresight and strategic intelligence, policies, structures and governance, government performance, global competitiveness, and innovation in government work. Officials also discussed performance indicators and how entities will avoid duplicating each other’s deployments. Two framings of the same project The workshop was guided by a stated principle: human leads, AI enables. Recalling how the project was introduced, when announcing the framework, Sheikh Mohammed bin Rashid Al Maktoum said the UAE would be “the first government in the world to largely deploy Agentic AI models” across its sectors and operations, and described systems that manage operations and run an independent series of actions without human intervention. The founding language is about autonomous execution. The implementation language is about human primacy. Both can be true across a portfolio of thousands of tasks, and the gap between them is exactly what the classification frameworks are meant to close. But until those frameworks are published, the question of who answers for an agent’s decision sits open. The material released so far covers pillars, training, indicators and governance structures. It does not describe what a citizen does when an autonomous system gets their case wrong, or where liability lands when it does. That is not a criticism unique to the UAE. No government has answered it, because no government has been here before. The difference is that the UAE is the one that will have to answer it first, at scale, on a two-year clock. The scale is the point Oversight sits with Sheikh Mansour bin Zayed Al Nahyan, with a taskforce chaired by Mohammad Al Gergawi, Minister of Cabinet Affairs. In May, the Cabinet approved the first package of government services to apply agentic AI tools, alongside the largest training programme in the history of the UAE government, covering 80,000 federal employees from ministers to new joiners. A national AI healthcare policy passed in the same session, as did frameworks making digital records the official source of core government data and requiring information to be collected once and shared securely across entities. The tempo since has been deliberate. A Dubai training event in June drew more than 300 participants from 50 federal institutions, with a stated goal of identifying which services and procedures suit agentic AI within 90 days. A separate workshop that month, run by the Presidential Court and the Ministry of Cabinet Affairs, put 600 staff through the same material. Why agentic AI in government travels beyond the Gulf Governments across Asia and Europe are approaching the same decision from a standing start, most of them with an AI strategy, a national body and no number attached to any of it. The UAE has attached a number and a date, which means that within two years there will be public evidence about what agentic AI in government actually does to service delivery, headcount and error rates. Nobody else has to run that experiment now. They do have to decide what they will do with the results, and whether they are willing to move before those results exist. The task classification frameworks are the document to watch. When they are published, they will be the first serious attempt by any state to write down which decisions a government is prepared to let a machine make on its behalf. Everything else in this programme is procurement and training. That part is constitutional. (Photo by Emirates News Agency) See also: OpenAI president urges enterprises to hasten AI security defences Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post Agentic AI in government just hit the hard part: deciding what a machine may decide appeared first on AI News. View the full article
  10. Advertising inside ChatGPT arrived with a promise that the assistant already knows what the user wants. So far, that hasn’t entirely been the case. Searchable, the AI visibility platform, analysed more than 11,000 ads served inside real ChatGPT conversations between 4 July and 4 August 2026, pairing every ad with the conversation it appeared in, then grading how closely the advertised product matched what the user was asking about. After assessing every ad according to one of three relevancy bands, a third (33%) of ads were found to be unrelated, with no connection to the conversation at all. Just over a quarter (27%) were a direct match, advertising the specific product the user was asking about. The largest group, 40%, were a contextual match, connected to something the user had raised earlier in the thread but not to the question in front of them. Two thirds of ads landed in conversations with no purchase intent anywhere inside them More than two-thirds (68%) of ads appeared in conversations where the user showed no sign of wanting to buy, book, hire or compare anything, at any point in the thread. Advertisers are seeing a large share of paid impressions appearing where no purchase was ever likely. Counted by conversation rather than by ad, 40% of ad-carrying chats contained at least one ad unrelated to the discussion, and in 28% every ad served was unrelated to the topic the user was discussing. In paid search, the advertiser picks the keywords that signal a buyer. ChatGPT’s ad platform works with context hints, where advertisers describe the conversations, situations and topics they want to appear in and OpenAI matches those against a live conversation. This takes place inside an AI engine that people also use to write, study, troubleshoot, and think through problems. The mismatch rate more than doubles between sectors Which sector an advertiser sits in predicts how badly its ads misfire. 47% of marketing and B2B services ads and 45% of software and SaaS ads were unrelated, the two worst rates of any sector measured. The most relevant targeting came from data brokers and background checks at 50%, travel at 43% and automotive at 42%. No sector managed a direct relevancy match for ads more than half the time. Insurance was rarely unrelated at 23% and rarely a direct match at 11%, because two thirds of its ads fall in the contextual band, attached to a life event somewhere in the thread rather than to the question being asked. OpenAI said its ads are becoming more relevant Denise Dresser, then OpenAI’s chief revenue officer, said in June that the rate at which users dismiss ads inside ChatGPT has fallen by half since the advertising business launched in February, and the company treats dismissals as its proxy for relevance. Searchable’s analysis covers a month roughly half a year into that rollout, and across it the share of ads with no connection to the conversation stayed close to a third, with no sign of improvement as ad volume grew. For brands, the consequence is that a placement inside an assistant is not a substitute for being the answer. An ad that lands in the wrong conversation is likely to be scrolled past, while the recommendation inside the response has already been shown to convert higher than other organic channels. The post A third of ChatGPT ads appear in irrelevant conversations appeared first on AI News. View the full article
  11. AI data centre regulation in Pennsylvania now begins with a signature. Before the state will so much as open a developer’s permit file, that developer has to sign a contract accepting a fixed set of conditions and the penalties for breaking them, and persuade the town that has to live with the building to say yes. Governor Josh Shapiro signed Executive Order 2026-05 in the US state’s capital, Harrisburg, on Tuesday. It directs the Department of Environmental Protection to review permit applications from proposed data centres only where the developer has committed to the Governor’s Responsible Infrastructure Development (GRID) Requirements through a Consent Order and Agreement, and has already secured local approval. It took effect immediately and covers all new applications. The template agreement was published the same day. No new law was needed. The order works through permitting authority the state already holds, which is why the rest of the industry should read it closely. Developers who refuse to sign are not banned. DEP simply will not open their file until every local approval is secured and every construction permit has already been reviewed and cleared, which inverts the sequence the sector plans around. The Department of Revenue is applying the same test to the state’s sales and use tax exemption for data centre equipment, and every data centre proposal has been removed from Pennsylvania’s permit fast-track programme, permanently. Under GRID, developers pay the full cost of the generation, transmission and distribution of their project needs rather than passing it to households and businesses. They must consult communities early enough to shape design decisions, hire and train locally, enter community benefit agreements, and meet water conservation standards. The secrecy clause is the part worth exporting One line in the order carries further than the rest. The use of nondisclosure agreements on data centre projects is no longer permissible for Commonwealth agencies. NDAs are the standard instrument in data centre siting worldwide. They are why residents in most host markets cannot find out a facility’s load, its water draw, or which hyperscaler is actually behind the shell company. Pennsylvania has decided that the secrecy is not a side effect of the deal but part of what went wrong with it. Two disclosure duties sit alongside. Operators must report annual energy and natural gas consumption, estimated average hourly use at peak, total water consumption and maximum day demand, plus any measures taken to protect the public from polluted air or water. The administration has also published a live map of every proposed project that has engaged with DEP, letting residents track each permit through review. Figures released with the order show why officials wanted the sorting done in public: more than 100 projects appear in public databases, 58 have approached DEP, 15 have applied for a permit, and five hold everything needed for a first phase. Data centres lose power before households do The order also sends the state’s Special Counsel for Energy Affordability to the Public Utility Commission with three asks. Utilities should recover the cost of PJM’s reliability backstop auctions from data centres rather than other customers. Data centre demand should be forecast accurately and disclosed. And when the grid comes under strain, data centres should lose electric service before anyone else. Curtailment ranking is the provision AI infrastructure planners will read twice, because it turns an argument about cost into a constraint on uptime. PUC chairman Steve DeFrank framed the principle simply: “growth should pay for growth.” Shapiro has travelled a long way to get here. He championed a US$20 billion Amazon commitment to sites in Bucks and Luzerne counties, and told developers this week that the world’s largest companies “can afford to be good neighbours, follow the rules, and do this right.” He faces re-election in November. The House passed bipartisan legislation codifying GRID; Senate Republicans declined to move it, and majority leader Joe Pittman said the governor’s remarks looked “aimed at erasing his remarks from last year.” Food & Water Watch, campaigning from the opposite direction, called the order “too little, too late,” pointing to the sales tax incentive that survives it. Why AI data centre regulation elsewhere now looks different Twenty-seven states in the US have been weighing large-load legislation, and California, Ohio and Utah have already enacted laws that go beyond the voluntary Ratepayer Protection Pledge developers signed at the White House in March, which carries no enforcement. New Jersey’s Data Centre Fair Share Act requires facilities above 50 MW to commit to at least 85% of projected power costs for a decade. Virginia added an energy consumption tax of $0.011 per kWh. Around 48 projects representing $156 billion in potential investment were blocked or stalled by local opposition in 2025. What makes Pennsylvania different is that it needed none of those instruments. No statute, no moratorium, no tax overhaul. A governor used existing permitting powers, attached penalties through a contract the developer signs voluntarily, and made the operating data public. Any administration facing the same constituent pressure can lift it, and the consent order template is already online. The variable to price into AI capacity is no longer land, power or ****** position. It is consent, and what a developer has to disclose to earn it. (Photo by Donnie Rosie) See also: OpenAI president urges enterprises to hasten AI security defences Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post AI data centre regulation just got a template that needs no new law appeared first on AI News. View the full article
  12. Autonomous artificial intelligence agents have already penetrated the offices of large, global enterprises. Now, HoneyBook is trying to bring that same capability to independent businesses with the launch of HoneyBook MCP, recently released as a connector for Anthropic’s AI assistant Claude. The move addresses a real gap. McKinsey’s recent State of AI survey found that nearly half of companies with more than $5 billion in revenue have moved AI beyond pilots and into everyday use across the business, compared with just 29% of companies under $100 million in revenue, leaving smaller businesses without the infrastructure that gives large enterprises an advantage. Founded in 2013, HoneyBook has developed a niche-leading CRM application that acts as a hub for SMEs, consolidating and managing everything from customers’ initial enquiries to their communications, bookings, contractual details, scheduling and payments. It’s a complete system of record that spans the entire customer relationship, and over time it becomes a veritable goldmine of data. However, far too many business owners have trouble making the most of this data, simply because they’re spread so thin managing operations. The reality of running a small business is that time tends to be a limited resource, and most operators can’t afford to spend hours each day digging through financial records, calendars and email conversations to surface those useful nuggets of information ahead of sales calls. This is what HoneyBook’s MCP connector aims to solve. Agentic AI value for SMEs With HoneyBook MCP, business owners can connect all of their business records directly to Claude (or any other AI service that supports MCP), making it far more accessible. Users can ask questions, such as what the status of a new lead is, or whether or not a client has paid an invoice yet, and Claude will immediately tell them. It can also take actions, helping users to update stages and dates based on customer interactions; send out invoices; create contracts from templates; deploy email reminders and so on. The launch of HoneyBook MCP is timely, because small business owners are eagerly embracing AI tools wherever possible. With all of the buzz and headlines around AI, these tools are impossible to ignore, and they’re having a profound impact. A recent HoneyBook study revealed that small business owners who’ve adopted AI earn almost five times as much revenue as those that haven’t. AI has the potential to provide a massive productivity boost, with 68% of small business owners saying that they believe advances in AI will be good for their business and 52% of AI-adopting small businesses reporting significant return on their investment, according to research from Bluevine. Changing the way business gets done HoneyBook MCP unlocks multiple ways for business owners to automate operational processes and save themselves valuable time. One of the most impactful examples is pipeline management. Busy owners who are dealing with dozens of clients at once can be excused for forgetting about one or two enquiries that might have gone stale. By the time they remember to follow up, that prospective client may have found another solution elsewhere. Claude can help prevent this. Simply ask which leads haven’t replied recently, and it will quickly pull up all of the relevant names so they can receive follow-ups. This kind of reconciliation was highlighted as one of the major benefits of HoneyBook MCP by early adopters of the solution. According to HoneyBook, pre-release preview users said it saved them hours that would otherwise be spent tracking unpaid invoices, surfacing old leads and fixing mislabelled business records. HoneyBook Senior Product Manager Helena Bachar said the ability to ask and receive answers about any aspect of their business provides massive value to HoneyBook users. “Giving our members a way to ask a business question and get the answer from what’s actually in their account, not a guess, changes how the workday starts,” she said. “The first thing people reached for wasn’t drafting or writing, it was their own track record: what got delivered, what got billed and what that tells them about the months ahead.” Solopreneurs likewise stand to benefit by using Claude as their access point to customer records. For example, photographers can quickly ask Claude which customers are still waiting for their photos and then instruct it to create a gallery and send them a link, as opposed to doing this manually. Once the job is done, they can then ask Claude to generate and send the invoice, or perhaps a standalone payment request for a one-off charge, without having to switch to a new tool. Service providers can use Claude to create a proposal and handle customer messages from directly within the chatbot’s interface. Instead of having to jump between different applications, users can simply instruct Claude to do much of the work for them, updating the necessary records along the way. It makes the entire client lifecycle visible inside Claude, transforming it from a useful chatbot into a personal assistant. HoneyBook also recognises the sensitive nature of business records, which is why the MCP runs each request within an isolated sandbox environment that’s expunged once the user’s session has finished, ensuring Claude never retains any private business data. It provides granular controls over what HoneyBook data Claude can access, too. Final thoughts For an industry built on personal relationships and hands-on service, the combination of real data access and real boundaries may be what finally turns AI from something small business owners are curious about into something they rely on every day. AI has a mixed track record when it comes to autonomous business operations, but with the right balance of thoughtful prompting and action, the potential is enormous. The post HoneyBook bets on agentic AI to streamline small business operations with its new Claude connector appeared first on AI News. View the full article
  13. OpenAI president and co-founder Greg Brockman warns that enterprise security teams face a compressed timeline to adopt AI defences. Brockman has published an account of what the company calls the “OpenAI-Hugging Face” incident, using it to argue that organisations need to uplevel their security practices with what he terms unprecedented speed. He writes that he has spoken with many organisations since the incident and found a consistent theme running through those conversations: leaders know they must move faster than their current security programmes allow. The urgency stems from a specific event. An “agentic collective” autonomously penetrated OpenAI’s own research infrastructure and then moved into the production infrastructure of Hugging Face. The attackers chained together previously unknown security flaws with leaked user account credentials found on the internet to complete the intrusion. Brockman calls it a preview of how a typical threat actor’s capabilities will evolve over the coming months. The AI defence decision facing security leaders Brockman argues the incident exposed a problem that extends beyond any single company’s network. He writes that accumulated technical debt inside every organisation “masks significant flaws” that defenders now need to locate and fix before attackers do. AI models developed across the industry are increasingly able to automate parts of real-world cyberattacks, he says, which makes long-standing security gaps easier to find and exploit. Those gaps range from bugs embedded deep in human-written software to forgotten permissions left unmanaged for years. The timeline for that decision is short by Brockman’s own account. Earlier in the year, OpenAI began releasing its cyber capabilities only to trusted defenders rather than the public, a deliberate attempt to keep defenders ahead. Since then, other companies have released open-weight models with cyber capabilities trailing the frontier by only a few months. Brockman points to a further model that appears scheduled for release at the end of August, which he says seems likely to accelerate the threat landscape significantly. For enterprise leaders, that compresses the window for building AI-assisted defences before broadly available models close the gap with attacker capability. Brockman frames the underlying dynamic as a race with two edges. AI-powered attackers will soon be able to find long-standing flaws across many existing systems, he writes, but the same technology gives defenders tools to find, prioritise, and fix those flaws faster. While describing security as remaining a cat-and-mouse game, Brockman argues that AI may shift its underlying economics in ways that favour defenders. OpenAI states it has begun training models specifically to write more secure code. Separately, the company points to its models’ capability in mathematical proofs, which it says can be applied to formally verify software security in ways that have proven difficult for human reviewers to achieve at scale. A test case against Brockman’s personal website Brockman offers a personal example of what faster response looks like in practice. After the incident, he asked ChatGPT Work, running publicly available GPT‑5.6 Sol, to assess the security of his personal site, gregbrockman.com. He describes it as a simple static site hosted on AWS with Cloudflare acting as a frontdoor, and says he expected limited surface area for vulnerabilities. The assessment took about 15 minutes and surfaced 13 issues. Brockman says many probably were not exploitable by themselves, but he could imagine them being chained together with other vulnerabilities. The tool found that his DNS records were not configured to prevent attackers forging emails from his address. His site was running an insecure version of jQuery and Cloudflare was forwarding requests to AWS over unencrypted HTTP. He then asked ChatGPT Work to fix the issues, which it did over roughly an hour. The tool opened the Cloudflare control panel in his browser and worked through DNS, TLS, and advanced security settings. It removed jQuery from the site entirely, migrated the site from AWS to Cloudflare Pages, and began a phased rollout of DMARC. Brockman says this as a small-scale demonstration of existing models operating as what he terms a cyberguardian, capable of finding a long tail of configuration issues that a human might lack the time or specific expertise to address, then applying fixes with an appropriately staged rollout. How OpenAI restructured its own defences Brockman writes that the Hugging Face incident showed OpenAI had underestimated the real-world cyber capabilities of its own AI models, prompting the company to strengthen its safety requirements and add urgency to existing safety research and internal security work. He sets out four areas of internal investment that inform his recommendations to other organisations. The first is using OpenAI’s own models to help secure its code. Codex, along with a security plugin, validates code changes and identifies vulnerabilities before deployment. Brockman is explicit that producing more findings requiring human validation is not the goal; the aim is catching real vulnerabilities before they ship and shortening the time between discovering an issue and deploying a fix. OpenAI’s ambition is to eliminate some classes of software vulnerabilities in newly-authored code. The second pillar involves using models to defend infrastructure on an ongoing basis. Brockman says almost all of OpenAI’s initial security alerts are now triaged by AI systems before humans get involved, which he says reduces workload for defenders and improves response time. The company is connecting these detections to bounded automated responses while keeping humans responsible for the highest-impact decisions, with the stated goal of detecting and responding to security issues at machine speed. Third, OpenAI uses its models to continuously enumerate and probe for potential attack paths, looking for vulnerabilities, misconfigurations, over-privileged identities, and unintended trust boundaries. This supports what Brockman calls ongoing assessment of the company’s security invariants, the properties it believes should hold true across its products and infrastructure. The fourth pillar is investment in fundamentals at scale, including secure architecture, defence in depth, and least privilege. The stated design goal is systems requiring multiple independent controls to fail simultaneously before anything catastrophic can occur. Network isolation, workload hardening, monitoring, and patching and deployment practices remain part of this baseline, and Brockman says they will matter more – not less – as AI capability increases on both sides. What Brockman tells enterprise security teams to do now Brockman sets out a list of actions for security teams, framed around speed rather than a full programme redesign. He recommends securing organisational buy-in and running tabletop exercises to model how these attacks might play out inside a given organisation. He advises giving security teams an agentic tool such as Codex or the Codex Security plugin, with approved access to codebases and infrastructure configuration, starting with the highest-priority systems rather than waiting for a company-wide rollout. He suggests equipping that agent with community-supported skills covering static analysis, security-focused code review, vulnerability variant analysis, and software supply-chain risk, then building organisation-specific skills around existing architecture and threat models. Organisations should run assessments against internet-facing services, authentication flows, infrastructure-as-code, and systems handling sensitive data first. Teams should then work through existing backlogs of scanner output, dependency alerts, and bug bounty reports, asking the agent to distinguish exploitable issues from noise. Brockman also recommends embedding agent-based review directly into development pipelines, checking for authentication mistakes, access-control bypasses, exposed credentials, and unsafe dependencies before code merges. For validated issues, he suggests having the agent generate a patch, write a regression test, and confirm the vulnerability no longer reproduces, while keeping human review for consequential changes. On automation, Brockman advises an incremental path rather than attempting to build an autonomous security operations centre immediately. Organisations should start with read-only scans of a single repository, move to advisory pull-request scanning, then live alert triage, and only later introduce automatic closure of narrowly defined false positives. A human should make every decision until confidence builds through that sequence. He also points organisations towards applying for Trusted Access for Cyber to gain approval to use GPT‑Daybreak‑Blue for defensive work including incident response, detection engineering, and malware analysis. Brockman recommends practising with the capability on logs and telemetry before an actual incident forces the issue. Brockman closes by arguing that no company can address this alone, calling on AI labs, security vendors, enterprises, and maintainers to share validated findings, fixes, and playbooks so that one organisation’s discovery strengthens the wider ecosystem. He describes the defender’s window as open now, with organisations needing to automate security programmes over the coming months to keep pace with attacker capability, ahead of the further open-weight model he expects at the end of August. See also: Alvys launches AI agents for freight TMS workflows Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post OpenAI president urges enterprises to hasten AI security defences appeared first on AI News. View the full article
  14. Freight software provider Alvys has launched an agentic AI platform that allows carriers and brokers to automate operational tasks directly within its transportation management system (TMS). Called Alvys Foundry, the platform supports pre-built and custom AI agents that work with freight data and workflows already held within the TMS. Foundry includes more than 20 pre-built agent templates. Customers can also create their own agents or work with Alvys engineers to build and configure agents around specific operating procedures. From AI assistance to execution The agents are designed to carry out defined steps in freight workflows, including detention processing, document handling, rate audits, shipment tracking, asset compliance, and claims management. Alvys’ Detention Agent can identify when detention time has been exceeded and file the detention. Document Intelligence reads and files rate confirmations, bills of lading, and proofs of delivery, while the Track & Trace Agent handles check calls and shipment status updates. Other templates include a Rate Audit Agent for checking invoices, an Asset Compliance agent for authority, insurance, and safety records, and a Claims Agent for opening and documenting claims. Customers do not necessarily have to program those workflows themselves. Alvys says operators can upload an existing standard operating procedure or describe a task in plain language, after which Foundry generates a workflow for approval. Agents can be tested against simulated data before deployment on live freight. Operators can then build, deploy, monitor, and pause them through the Foundry platform. Alvys CEO Nick Darman said the platform was developed after the company observed how customers were adopting AI across their freight operations. He said some implementations introduced additional tools and logins without fitting into existing operating procedures. Alvys already uses AI elsewhere in its TMS. Insights provides information across loads, lanes, and margins, Intel issues alerts on issues including weather and cargo theft, and another AI feature creates loads from uploaded rate confirmations. Foundry adds agents that can perform approved tasks within configured workflows. Foundry operates on the existing infrastructure behind Alvys’ TMS, which includes more than 120 integrations and native electronic data interchange connections with hundreds of shippers. This gives the agents access to operating context such as lane history, customer rules, documents, margins, appointments, and exceptions. “We have the freight context and we understand your lanes,” Darman said. He added that keeping the agents within the existing platform removes the need to maintain separate integrations and logins. Alvys said Foundry runs on a SOC 2-compliant security foundation and that its agreements with model providers prevent customer data from being used to train public models. Foundry’s Agent Shield governance layer lets operators set approval and spending thresholds for individual agents. Higher-impact actions can be paused for human sign-off, while agent decisions, actions, and manual overrides are recorded in an audit trail. Foundry also includes a model-selection system that can route tasks between large language models based on factors including cost, speed, and quality. The system is designed to manage computing costs and avoid tying workflows to a single model provider. AI agents expand across freight workflows C.H. Robinson is also using AI agents for operational freight tasks. The freight broker told DC Velocity in January that it had deployed more than 30 agents that had collectively completed millions of tasks previously handled manually. “Agentic AI doesn’t just analyse or generate content; it acts autonomously to achieve goals like a human would,” Mark Albrecht, vice president of artificial intelligence at C.H. Robinson, said. One C.H. Robinson agent processes more than 10,000 emailed pricing requests a day by reading the request, obtaining a price from the company’s pricing system, and sending a response. Another reads load tenders and attachments before converting the information into orders. Uber Freight has also embedded AI agents into its transportation management software. In May 2025, the company said it had more than 30 agents automating work across procurement, shipment execution, tracking, payments, and analytics. Uber Freight said it wants its TMS to move beyond recording logistics information by automating repetitive operational work and guiding users through transportation processes. The company described its longer-term goal as moving the TMS beyond a system of record and toward a platform that can proactively guide users while automating repetitive tasks. Alvys is applying that execution model within its own TMS rather than through a separate automation layer. Customers can start with pre-built agents, have Alvys engineers configure custom agents, or build their own workflows on the Foundry platform. Spartan Carrier Group is among Foundry’s early customers Founder and CEO Carlos M. Llanes Jr. said the platform is being used to reduce manual freight work while employees remain focused on judgment and service. The launch follows Alvys’ $40 million Series B funding round in September 2025, led by RTP Global. FreightWaves reported that the company has raised $77 million in total, while Alvys says its platform handles more than $9 billion in freight annually. Foundry is being introduced through customer cohorts rather than an unrestricted general rollout. FreightWaves reported that the first cohort filled after Alvys presented the platform at a June 17 customer advisory board meeting, while a waitlist is open for the second cohort. (Photo by Barrett Ward) See also: Hershey applies AI across its supply chain operations Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events, click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post Alvys launches AI agents for freight TMS workflows appeared first on AI News. View the full article
  15. Zhipu’s release note for GLM-5.3 contains a sentence that did not make it into most of the coverage. Describing its own cybersecurity results, the Beijing company writes that capability “is growing fastest exactly where we are furthest behind.” Zhipu, which also trades as Z.ai, is one of a handful of ******** labs releasing models that compete with the American frontier. On August 14, it launched GLM-5.3, a coding-focused model, and published a technical release note setting out how the model performs against its rivals. That note is the source for everything reported here. The claim that travelled was about security. Alongside the coding results, Zhipu said GLM-5.3 had become unexpectedly good at finding software vulnerabilities, scoring 84.5% on a benchmark called CyberGym against 83.8% for Anthropic’s Mythos 5 and 83.6% for OpenAI’s GPT-5.6 Sol. Headlines followed reporting that a ******** model now out-finds the American ones at bug hunting. The reason that lands harder than a usual benchmark result is what vulnerability discovery has become. A model that can read a codebase and locate exploitable flaws is useful to a defender auditing their own software and useful to anyone doing the same to somebody else’s. Anthropic’s equivalent work sits behind restricted access for that reason, while Zhipu intends to publish GLM-5.3’s weights for anyone to download. Zhipu’s own release is more measured than the coverage it produced. The CyberGym number is real, and it is in the paper. It is also the narrowest of the three cybersecurity results the company published, and Zhipu is upfront that the other two go the other way. Three benchmarks, three different pictures CyberGym starts from source code the model can read and tests whether it can find a vulnerability and confirm the flaw is genuine. That is the result that travelled, and the margin is seven tenths of a percentage point. ExploitBench asks something harder, requiring the model to reason about a real vulnerability and how it would be exploited. GLM-5.3 scores 54.4%, more than double its predecessor’s 24.4%. Mythos 5 scores 78.0% and GPT-5.6 Sol 76.5%. ExploitGym counts how many exploitation tasks a model finishes inside a fixed time budget. GLM-5.3 completes 105 tasks in two hours and 130 in six. Mythos 5 completes 181 and 247. Those two results have been reported thinly, and they are the ones that describe the gap. Finding a flaw and building a working exploit from it are different jobs. Zhipu’s reading is that the further along that chain a test sits, the further behind its model is, and the company says so in the release rather than leaving it to be discovered. Zhipu’s own comparison across the three cybersecurity benchmarks. GLM-5.3 leads on CyberGym, the vulnerability discovery test, and falls behind Anthropic’s Mythos 5 on both exploitation measures. Source: Z.ai. Which Anthropic model, and why it keeps changing Part of the confusion in the coverage comes from Zhipu comparing three different Anthropic models in three different places. The main benchmark table sets GLM-5.3 against Opus 4.8. The performance charts use Fable 5. The cybersecurity section uses Mythos 5. Anyone reading quickly comes away with a single comparison that does not exist. On coding, the picture is mixed rather than dominant. GLM-5.3 leads Opus 4.8 on some tests and trails it on others, and Zhipu states plainly that its model remains behind Claude Fable 5 on the company’s own internal coding benchmark. How the tests were run The methodology footnotes contain something the summaries skipped. Zhipu evaluated GLM-5.3 on CyberGym, ExploitGym, ExploitBench, Terminal Bench and several other tasks inside Claude Code 2.1.207, Anthropic’s coding agent. That is not improper. Using a common harness across models is how a comparison stays fair, and Zhipu documents the settings it used. It is worth noticing anyway. A ******** open-weights model’s frontier claims are being measured through American agent software, which says something about where the tooling layer sits in this competition that the model scores do not. Two further details deserve attention before the CyberGym result is treated as settled. The score is a single run, reported as pass@1 across 1,507 tasks, with no variance figures given. A gap of seven tenths of a point between two single runs is not a gap anyone should lean on. And the ExploitGym time budgets were normalised using throughput rates from Artificial Analysis, with rescaling factors listed for GLM-5.3, Kimi K3 and Qwen3.8 Max, but not for Mythos 5. The vulnerability count and the number that is missing Beyond the benchmarks, Zhipu says it worked with security teams in China to run its models against real codebases, identifying 2,436 vulnerabilities across 269 open-source projects. The severity split is 107 critical, 990 high, 1,286 medium and 53 low. The oldest flaw dates to 1981, and the average vulnerability had been sitting in code for 26.6 years before it was found. Zhipu’s disclosure summary. The panel labels 1,097 findings as critical and high, matching the severity breakdown of 107 critical and 990 high. The body text of the same release describes the figure as medium-to-high. Source: Z.ai. One discrepancy is worth carrying carefully. Zhipu’s summary panel labels 1,097 findings as critical and high, which matches the severity table. The body text of the same release describes those 1,097 as medium-to-high. Several outlets have reproduced the second version. The count also arrives after what Zhipu describes as expert review, screening and deduplication, so the raw model output is not what is being reported. Of the 2,436 findings, 53 have been publicly disclosed, and 2,383 remain under embargo. The release does not say how many were previously unknown, and it does not say how many were independently reproduced. Those are the two figures that would turn a volume claim into a capability claim. What matters more than the benchmark table Two things in the release have longer consequences than the CyberGym margin. The first is efficiency. Zhipu reports GLM-5.3 reaching 31.4% on its internal coding benchmark at around 50,000 output tokens per task, against Opus 4.8 at 29.5% using 120,000. Slightly better work for less than half the tokens is a cost argument, and cost determines whether security teams outside the largest budgets can run these tools at all. The second is distribution. Zhipu says the weights will be published once safety evaluation and hardening are finished. That has not happened yet, and until it does, the open-weights claim is a commitment rather than a fact. If it holds, a model with documented vulnerability-discovery capability becomes something any team can download and run locally, including in markets that will never have access to an export-controlled American model. The weights are due at the end of August. See also: Anthropic walks into the White House and Mythos is the reason Washington let it in Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post Reading Zhipu’s GLM-5.3 results past the headline number appeared first on AI News. View the full article
  16. Samsung Research America’s Digital Health Team has presented two AI foundation models designed to learn from wearable biosignals. The work centres on data captured by smartwatches, including heart activity, sleep, and physical activity. The company discussed its Connected Care vision at the Health Forum during Galaxy Unpacked in July 2026. Samsung described a future of preventive, personalised, and connected care, supported by health technology and healthcare partnerships. Its research team positions health foundation models as one component of new consumer health experiences. Sharanya Desai, Head of Digital Health Algorithms at Samsung Research America, said: “This research is significant because it lays the technical groundwork for delivering health insights that are efficient, precise, and continuous through a health foundation model. “We will continue to develop and advance health foundation models that can be applied to a variety of biosignals and health features that can operate on-device with limited sensors and computing resources.” Samsung’s health AI foundation model research A health foundation model uses self-supervised learning to identify features in unlabeled biosignal data. Samsung says that pretraining on large health datasets allows one model to support downstream tasks such as biosignal analysis, biomarker development, and health issue prediction. The research covers two models with different aims. xMAE, short for Physiology-Aware Masked Cross-Modal Reconstruction for Biosignal Representation Learning, learns temporal relationships between different biosignals. HiMAE, or Hierarchical Masked Autoencoder, learns health patterns across multiple time scales in wearable time-series data. Samsung says xMAE was accepted to the International Conference on Machine Learning. HiMAE was accepted to the International Conference on Learning Representations. The company describes both as work on physiological relationships and temporal structures in biosignal data. The models address different parts of wearable-data analysis. xMAE connects two cardiac signals that measure related activity through different mechanisms. HiMAE analyses data at short and long intervals, allowing one pretrained model to support classification, numerical prediction, and data generation. xMAE links continuous PPG data to ECG signals Electrocardiograms, or ECGs, measure the heart’s electrical activity directly. Samsung describes ECG as useful for measuring heart rate and heart-rate variability. It can also identify abnormal heart rhythms and risks associated with conditions such as atrial fibrillation. Wearable ECG readings generally require a user to pause and take an active measurement. Photoplethysmography, or PPG, takes a different approach. PPG detects changes in blood flow and can run passively through sensors in wearable devices such as smartwatches. Both signals originate from cardiac activity. They occur with a time difference, which Samsung compares with hearing thunder after seeing lightning. xMAE learns that temporal relationship by reconstructing masked parts of an ECG signal from PPG data. This design aims to analyse cardiovascular-health features through continuously measured PPG data without separate manual ECG measurements. The model’s pretraining used about 9,400 hours of ECG and PPG data. Subbu Venkatraman, Head of the Digital Health Research Lab at Samsung Research America, commented: “Biosignals are inherently dynamic, with unique time-varying physiological properties. The key contribution of this research lies in proving the viability of health foundation models capable of capturing both the inter-signal relationships and their underlying temporal structures. “We remain committed to advancing foundational health AI research and translating it into healthcare solutions that meaningfully improve people’s health and wellbeing.” Samsung reports that xMAE outperformed unimodal biosignal models and existing multimodal learning methods in 15 of 19 evaluation tasks. Those tasks covered cardiovascular disease prediction, abnormal test-result detection, and sleep-stage classification. The company also says the learned features showed potential for use across sensor devices, body locations, and data-gathering environments. HiMAE analyses wearable data across time scales Wearable data can carry different information over different time periods. Short segments can show fast-changing signals such as heartbeats. Longer segments can reveal patterns that build over time, such as sleep or physical activity. HiMAE uses multiple encoders to analyse short and long data segments separately. Samsung says this arrangement enables the model to identify the time scale needed for a health task. Heart-rate analysis and sleep prediction can therefore draw on different parts of the time-series data. The training method reconstructs masked portions of wearable data. Samsung says this lets HiMAE learn patterns from biosignals where labelled data is limited. The model then supports classification, numerical prediction, and data generation from a single pretrained system. Samsung says HiMAE achieved high performance with a smaller model than existing models. The company also reports that it can produce results in less than one millisecond on a smartwatch-class central processing unit. That processing claim places the model’s analysis on the device rather than on cloud servers. Foundation models trained on unlabelled physiological streams provide a mechanism to extract diagnostic markers, run predictive health classifications, and generate user guidance from consumer hardware all without continuous server connectivity. See also: Google AI health coach to use Abbott glucose data Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post Samsung health AI models analyse wearable biosignal data appeared first on AI News. View the full article
  17. Abbott and Google are linking continuous glucose monitoring data with Google’s AI-powered health coaching tools, giving the Gemini-powered service access to another source of personal health information. Under a multiyear agreement, data from Abbott’s Lingo continuous glucose monitor will be integrated into the Google Health app. Users will be able to view glucose trends alongside information related to activity, sleep, and other wellness metrics. Google Health Coach, which is built with Gemini, will use Lingo data alongside other health information to provide recommendations covering nutrition, activity, sleep, and recovery. Access to the AI coach requires a Google Health Premium subscription. Google says Health Coach is not intended for medical purposes and warns that its AI responses can be inaccurate or incomplete. The company advises users to verify its responses and consult healthcare professionals when appropriate. Google had already been expanding the types of health information available to Health Coach. The company said in May that US users could sync medical records with the Google Health app, including laboratory results, vital signs, and medication information, and share those records with the coach to ask questions or receive summaries. The Google Health app can also receive information from compatible wearables and third-party applications through Health Connect, Apple Health, and Google Health APIs. The companies expect the Lingo integration to begin rolling out in the Google Health app later this year. AI coaching meets continuous glucose data Lingo is an over-the-counter continuous glucose monitoring system designed for adults aged 18 and older who do not use insulin. The wearable tracks glucose levels throughout the day, while its accompanying app shows how factors including food, movement, and other daily activities correspond with changes in glucose. According to FDA documents, Lingo continuously measures glucose in interstitial fluid rather than directly from blood. Its sensor is inserted under the skin on the back of the upper arm and uses an electrochemical process to measure glucose before transmitting the readings to the Lingo app through Bluetooth Low Energy. The FDA says the Lingo app can display real-time glucose values, trend arrows, and glucose graphs. The sensor can be worn for up to 14 days, but the Lingo app does not provide glucose or system alerts. The product is based on technology used in Abbott’s FreeStyle Libre platform. Abbott made Lingo available without a prescription in the US in 2024, and the system is also available in the ***. Although Lingo shares underlying sensor technology with FreeStyle Libre 2, the products have different intended uses. FreeStyle Libre 2 is cleared for diabetes management, while the FDA describes Lingo as a system for helping adults who do not use insulin understand how glucose readings relate to nutrition, exercise, and daily activities. FDA documentation also states that users should not take medical action based on Lingo readings without consulting a qualified healthcare professional. Abbott and Google are also planning a large real-world study focused on metabolic health. The study will combine continuous glucose readings with wearable, laboratory, and survey data to examine relationships between activity, sleep, well-being, and metabolic health. Abbott and Google said the findings will be used to refine Google Health Coach and inform future Lingo features. Google connects more health data to Gemini The Google Health app brings information from multiple sources into one place. Users can sync activity, fitness, sleep, vital-sign, and medical-record data and connect compatible applications and devices. Google says users can decide what information they save, turn optional features on or off, export their data, and delete it. The company also states that Google Health data is not used for Google Ads. Google Health Coach can use personal health records to personalise its responses, while the Google Health app can sync information from compatible apps and devices. The Abbott integration will add ongoing glucose data from Lingo. Google has also been extending Gemini into other healthcare-related tasks. Zocdoc announced this week that US users can search for healthcare appointments and book providers directly through the Gemini app. The Zocdoc connected app provides real-time appointment availability from a network of more than 200,000 providers across more than 200 specialities. Zocdoc said it is Gemini’s first connected-app partner for health appointments. (Photo by Towfiqu barbhuiya) See also: Novo Nordisk and AWS bring agentic AI into drug discovery Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post Google AI health coach to use Abbott glucose data appeared first on AI News. View the full article
  18. Okta says identity-scoped Model Context Protocol (MCP) tool lists can reduce AI agent token costs. Each model call made by an AI agent can include schemas, names, descriptions and parameters for every tool exposed by a MCP server. Okta calls the resulting prompt overhead the “tool tax”: tokens consumed as a model considers tools, including those it will never call. The company argues that this cost appears before an agent attempts a tool call. A later rejection of an unauthorised request therefore cannot recover prompt tokens already consumed. Okta’s proposed control filters the list of tools before it reaches the model, using permissions assigned to an agent identity and the user associated with it. Okta’s internal modelling found that some permission scenarios reduced the number of visible tools by more than 90%. The company said tool-schema costs fell by roughly the same proportion, although it did not provide absolute token or dollar figures. MCP tool schemas create prompt overhead on every turn MCP servers have become a route for connecting AI agents to tools and data. Okta cites connections to Google Workspace, Slack and internal MCP servers as examples. An MCP server can expose a large number of tools, and the model receives a representation of each available tool in its prompt on every turn. That representation includes a schema. It also includes the tool name, description and parameters. Okta says the cost compounds when a widely used MCP server exposes many tools. Each active user incurs the prompt overhead whenever their agent makes a model call. The company frames this as both a tool-count problem and a user-count problem. The issue also has an access-control dimension. An agent that sees tools outside its authorisation scope can attempt to use them. A control that rejects the call at runtime can block execution, though the model has already received the tool definition and used tokens to process it. Okta filters tools before the agent prompt is built Okta positions the capability within its “blueprint for the secure agentic enterprise”, which asks organisations to identify their agents, their permitted connections and their authorised actions. Its approach narrows the connection question from access to a whole MCP server to access to individual tools on that server. An administrator configures the tools that a particular identity may use in the Okta dashboard. Okta then returns the scoped tool set instead of the server’s full catalogue. The agent receives this shorter list in its prompt for each turn. Okta says it checks scope again at runtime before a tool call executes. This design applies least-privilege access at the tool level. The company says an agent should not be aware of resources, databases or tools that it has not been expressly authorised to use. Removing unavailable tools from the prompt also removes their schema cost from the model call. Okta does not describe a live customer deployment in the post. Its evidence for the claimed reduction comes from internal modelling using Okta product data and public vendor documentation, with no customer data used. Internal model used OAuth scopes and representative roles Okta modelled a single MCP client with access to a catalogue of enterprise tools. It compared the number of tools visible to the model before and after identity-based scoping. To estimate scoped exposure, the company mapped Okta MCP Server tools to the OAuth scopes that unlock them. It then defined representative user segments. These included helpdesk read-only users and helpdesk operators. Other segments were app administrators, brand and email administrators, and super administrators. Okta weighted each segment according to an assumed share of monthly traffic. The company calculated tool-count reduction as one minus the ratio of scoped tools to unscoped tools. It said some scenarios removed more than 90% of visible tools. Its post states that tool-schema token cost tracks tool count nearly linearly because each tool contributes its name, description and parameter schema to every prompt. Okta says actual results vary according to the tool catalogue, distribution of permissions and model selected. Average schema size, request volume and model pricing also affect absolute token and dollar costs. Okta contrasts identity entitlements with gateway spending controls The post distinguishes identity-based scoping from gateway controls. Okta says gateways can cap spending by key, team or group, and can support routing and rate limiting. A gateway can meter tokens entering and leaving a system, as well as dollars spent. Okta says those controls can limit costs after a model decision becomes expensive. Identity entitlements provide a different input. Okta says per-user and per-agent entitlements can determine the tools available to a specific agent or the person behind that agent, rather than applying access information at group level. Paul Webber, Principal Cybersecurity Industry Analyst at Software Analyst Cyber Research, said: “Cost control for agents is best provided using identity governance tools that offer more granular control and precision without disrupting business processes. “Okta’s approach is an elegant way to do this because it leverages the same entitlement data that governs security, not a separate metering layer without that insight.” Okta’s account presents the gateway as a control for what passes through it. The identity layer filters the available tool set before those tools need to be metered. Tool visibility also affects MCP attack exposure The post ties the same mechanism to security exposure. Okta says removing tools from an unauthorised identity’s view also removes actions that identity could take if it were compromised. Its proposed scope check operates at two points. The first occurs as the tool list is assembled for the agent prompt. The second occurs when the agent attempts to execute a tool call. Okta describes the result as a smaller blast radius for a compromised identity. The remaining exposed tools determine the set of actions available to that identity. In the company’s model, the prompt contains only tools associated with the identity’s authorised OAuth scopes. For organisations assessing MCP access, tool inventory and entitlement mapping are the main operational inputs. Okta’s methodology maps MCP Server tools to the OAuth scopes that unlock them, then compares the full tool catalogue with the scoped catalogue visible to each representative user segment. Okta is a key sponsor of this year’s AI & Big Data Expo Europe held in Amsterdam on 19-20 October 2026. See also: Meta Muse Glimmer brings local AI agents to consumer GPUs Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post Okta targets AI agent token costs with MCP scoping appeared first on AI News. View the full article
  19. Google’s research medical AI system, AMIE (Video), conducted synchronous video consultations with professional patient actors and received clinical evaluator ratings on par with primary care physicians across several core measures. Fifteen trained actors portrayed conditions across cardiopulmonary, abdominal, HEENT, neurological or psychiatric, and musculoskeletal presentations. Google says studies involving real patients and their own health conditions must follow before anyone can draw conclusions about clinical use. AMIE divides a video consultation among three agents AMIE uses an asynchronous multi-agent architecture rather than assigning dialogue, clinical reasoning, and perception to one model process. Google says a single agent cannot currently sustain natural conversational response times while also conducting detailed reasoning and continuously processing audio-visual input. The talker agent handles the spoken interaction with the patient. It aims to maintain conversational flow, drawing on information from the other agents. The planner agent runs in the background, updating differential diagnoses and management plans as the consultation progresses. It also identifies missing information and reprioritises clinical goals. A perception agent reviews video and audio streams continuously. It looks for non-verbal signs, physical findings, and auditory signals, then places those observations into the conversation’s clinical context. Latency remains central. Deep clinical reasoning takes time, and long pauses can affect rapport during a consultation. Google’s architecture separates the patient-facing dialogue from the slower work of reasoning and perception, allowing the talker agent to respond without waiting for every background process to finish. Google reports that automated evaluations found each agent improved clinical measures, including history-taking, clinical reasoning, and treatment recommendations. The evaluations also covered patient-centred communication and response latency. The study compared video AMIE with text and physicians Google structured the human evaluation as a multi-arm randomised study. AMIE completed real-time video consultations. A text-only AMIE version provided a modality baseline. Ten board-certified primary care physicians used the same video interface. An independent panel of 20 experienced primary care physicians reviewed every consultation using established clinical rubrics. The panel assessed general clinical competence, then applied scenario-specific criteria tailored to the case. The study covered five body systems. Each scenario followed a standardised consultation format with a trained patient actor. Google says evaluators rated AMIE on par with the PCP group for history-taking thoroughness, diagnostic accuracy, management appropriateness, and communication quality. AMIE (Video) also matched or exceeded AMIE (Text) across those measures. Evaluators rated the video system higher on eliciting physical signs and proactively guiding actors through virtual examination manoeuvres than either the PCP group or text-only AMIE. Case-specific perception and examination scores reflected the same reported pattern. Patient actors also preferred the synchronous video interface to text chat. Google says they rated video as easier to use and more effective for communicating health concerns. The actors rated AMIE favourably for empathy, rapport, and confidence in care when compared with both study alternatives. Automated testing came before physician review Google built an automated evaluation suite to develop the video system before the human study. Its framework draws on a taxonomy of telehealth competencies from medical literature, covering visual cues, auditory signals, and physical examination manoeuvres. Single-turn assessments tested specific perception and reasoning tasks. Google gives anatomical laterality and signs of respiratory distress as examples. Multi-turn simulated audio consultations assessed the system’s conversational performance over a longer interaction. Visual input entered some of these multi-turn simulations as text descriptions. In a Parkinson’s scenario, for example, an AI patient simulator could describe a patient holding paper to the camera showing cramped, tiny handwriting. This arrangement helped Google test dialogue behaviour alongside visual information, though it does not replicate an end-to-end live video feed. The automated suite allowed rapid changes to the system design and exposed capability gaps before the actor-based study. The subsequent OSCE evaluation used a synchronous video consultation interface, although the patient presentations still followed prepared scenarios. The split between these methods should shape procurement and governance discussions. Automated assessments can test defined perceptual tasks at scale. Simulated video consultations can assess interaction quality under controlled conditions. Neither method establishes performance with patients whose symptoms, behaviour, connectivity, environment, and medical history fall outside a prepared case. Production evidence remains limited to text-based work Google identifies several limits in the AMIE research. Professional actors cannot fully reproduce the variability of real patient encounters. The scenarios also excluded presentations that actors could not portray authentically, including cases where audio-visual perception may carry more diagnostic value. Targeted automated evaluations found occasional perception and reasoning errors. Google also reports intermittent technical issues that can interrupt conversational naturalness. Project Astra remains a prototype, with system-level technical considerations outside this medical application. Google states that real patient research is the next stage. The company has begun related work in clinical settings with the text-based AMIE. A feasibility study with Beth Israel Deaconess Medical Center provided initial evidence on safety and utility in clinical practice, Google says. An ongoing nationwide randomised study with Included Health is evaluating AI in real-world virtual care. The Google study provides controlled evidence on video consultation behaviour, physical-examination guidance, and clinician scoring. It does not yet provide evidence that AMIE can safely diagnose or manage real patients in production. See also: Novo Nordisk and AWS bring agentic AI into drug discovery Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post Google tests AMIE for clinical video consultations appeared first on AI News. View the full article
  20. Novo Nordisk is expanding its use of AWS artificial intelligence tools across drug discovery, including AI agents for target identification, therapy design, and research workflows. Under the agreement announced recently, AWS will become Novo Nordisk’s preferred cloud provider and strategic AI partner. The companies have also created a co-innovation hub at Novo Nordisk’s existing London facility, where AWS engineers and Novo Nordisk scientists will work directly with the pharmaceutical company’s data and research insights. The London hub will bring Novo Nordisk R&D staff together with AWS engineers, AI specialists, applied scientists, and professional services teams. AWS said the arrangement is intended to reduce handoffs between computational analysis and laboratory research as potential drug candidates move through early development. Thilde Hummel Bøgebjerg, executive vice president of Enterprise IT & Quality at Novo Nordisk, said the partnership combines AI technology with the company’s scientific expertise. “AI has the potential to transform how we discover, develop and deliver medicines, but real impact comes from combining advanced technology with deep scientific expertise and a clear focus on patients,” Bøgebjerg said. The companies said the hub is intended to shorten the path from identifying a drug target to a first human dose. The partnership will combine Novo Nordisk’s research in chronic diseases with AWS cloud infrastructure, AI services, and life sciences technologies. AI enters more research workflows Novo Nordisk already uses AWS technology across internal AI workloads. According to an AWS case study, more than 25,000 Novo Nordisk employees have used a generative AI platform built on Amazon Bedrock to create chatbots covering more than 2,500 use cases. Those applications include information retrieval and document drafting across non-regulated processes. One of the larger deployments uses a collection of about 140,000 documents and processes more than 26,000 prompts each month, according to AWS. Novo Nordisk has also applied generative AI to clinical-study documentation. AWS said a system using Anthropic’s Claude 3.5 through Amazon Bedrock reduced the time required to generate some clinical documentation by more than 90%. According to the AWS case study, work that previously involved 40 to 50 people and could take as long as 15 weeks was reduced to minutes for a team of three. Medical professionals continue to review and validate the generated material. The new agreement moves that use of generative AI further into drug research. One of the technologies involved is Amazon Bio Discovery, which is designed to support computational biology and drug development using AI models and agents. AWS said Bio Discovery provides researchers with access to more than 40 biological AI models. AI agents can be used to select and coordinate models for different research tasks, while organisations can also combine AWS-hosted models with their own proprietary models in multi-step workflows. Novo Nordisk plans to use the service for tasks including identifying potential drug targets, designing therapies, and analysing biological information. The platform can also generate and rank potential drug candidates before selected candidates are sent for laboratory synthesis and testing. Experimental results can then be returned to the computational workflow for further analysis and model refinement. This links computational drug design with physical testing in the same research workflow. Novo Nordisk will also use Amazon Bedrock to build AI applications that can work across clinical, genomic, and imaging datasets. According to the companies, connecting these data sources is intended to allow findings from early-stage research to inform clinical trial design. Novo expands use of AI agents Another part of the agreement involves Amazon Bedrock AgentCore, AWS infrastructure for deploying and operating AI agents. AWS said AgentCore allows agents to work across enterprise data, connect with existing systems, and carry out multi-step workflows. Novo Nordisk has not detailed which AgentCore functions it plans to use in individual research workflows. The company said it intends to deploy the technology across research and operational processes. Dan Sheeran, vice president and general manager of Healthcare and Life Sciences at AWS, said the partnership is focused on applying agentic AI across drug discovery. “Novo Nordisk’s partnership with AWS shows what’s possible when you pair AI with deep life-sciences expertise — removing bottlenecks across drug discovery, not just studying them,” Sheeran said. Novo Nordisk’s earlier Bedrock deployments were focused mainly on employee-facing generative AI applications. The new agreement extends its use of AWS AI tools into scientific and operational workflows linked to drug development. AWS CEO Matt Garman said the companies intend to use the technology to shorten drug discovery timelines, process complex research datasets, and apply findings across therapeutic areas. AWS Forward Deployed Engineers will also work directly with Novo Nordisk teams on the systems developed through the partnership. The programme embeds AWS technical staff with customers to work on AI projects, including agent-based applications. The agreement builds on a broader relationship between Novo Nordisk and Amazon that also includes Amazon Pharmacy, Amazon Ads, and One Medical. AWS has separately announced life sciences AI collaborations with Flagship Pioneering and Johnson & Johnson in 2026. Novo Nordisk has also pursued AI projects outside AWS. The company has worked with OpenAI on deploying artificial intelligence across areas including drug discovery, manufacturing, and commercial operations. It has also signed on to use Denmark’s Gefion supercomputer. The system provides computing capacity for workloads including AI model development and other data-intensive scientific research. Novo Nordisk CEO Mike Doustdar said the AWS partnership is intended to accelerate drug discovery and expand the use of AI across the company’s work on chronic diseases. (Photo by Google DeepMind) See also: Why health AI interfaces must adapt to user expertise Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post Novo Nordisk and AWS bring agentic AI into drug discovery appeared first on AI News. View the full article
  21. Artificial intelligence is giving security researchers new ways to examine code, trace unusual behaviour and identify flaws that conventional tools may overlook. The pressure is particularly visible around Zero-day vulnerabilities, a recent Minimus analysis considers how container composition, dependency records and rebuild speed affect the response after an unknown flaw is exposed. Faster analysis helps only when organisations can also establish where the vulnerable software is running. AI is finding flaws that traditional tools may miss In May 2026, Google Threat Intelligence Group reported the first case in which it believed a threat actor had used AI to help develop a zero-day exploit. The exploit appeared in a Python script and bypassed two-factor authentication on a widely used open-source system administration tool when valid credentials were already available. Researchers said they had high confidence that an AI model assisted with both discovery and weaponization. Their assessment drew on the script’s unusually detailed instructional comments, a fabricated vulnerability score and a highly structured coding style associated with generated output. Google did not claim that the wider operation was autonomous or attribute the code to a particular model. The flaw itself is what makes the case significant. It involved a hard-coded trust assumption rather than a ******, memory error, or unsafe input. Fuzzers and static-analysis tools are well suited to finding many conventional implementation problems. A language model can also examine how permissions, functions and expected behavior interact across a codebase. That creates another route to finding logical contradictions that leave no obvious technical trace. Google’s wider data suggests this was not an isolated concern. According to Google Threat Intelligence Group’s 2025 analysis, researchers tracked 90 zero-days exploited in the wild during 2025, compared with 78 in 2024. Enterprise software and appliances accounted for 43 cases, or 48% of the total. Both figures were records in Google’s dataset. Complex containers make exposure harder to trace Once a flaw becomes public, security teams first have to work out where it is running. That can be difficult inside a container environment. An image may contain operating-system packages, application libraries and dependencies inherited from its base image, alongside shells or utilities with little connection to the workload’s visible purpose. A vulnerable component can therefore sit several layers below the application itself. It may appear across numerous images even when the organisation never added it directly. Log4Shell exposed this problem at scale in 2021. The affected Log4j library had been incorporated into a wide range of products and services. For many organisations, obtaining the patch was only the beginning. They still had to identify every server, application and container carrying a vulnerable version before they could complete remediation. Software bills of materials provide a clearer record of what each image contains. Smaller images can also reduce the search by excluding packages that the workload does not need. Minimus examines the issue through package reduction, dependency visibility and the rebuilding of images after an affected component is disclosed. The benefit is simpler than preventing zero-days altogether. A minimal image can still contain an unknown flaw. It gives teams fewer packages to investigate, fewer possible exposure points and less software to replace or retest once the problem becomes known. AI-generated fixes still need software context AI is also being used to shorten the time between disclosure and patch development. Models can inspect source code, compare vulnerability reports with package records and propose changes for affected versions. None of that is especially useful when package records are outdated or nobody knows which images contain the vulnerable component. Earlier coverage of an AI agent designed to automate vulnerability fixes detailed how Google DeepMind’s CodeMender contributed 72 security fixes to established open-source projects during its first six months. The system combines model reasoning with static analysis, runtime testing and fuzzing to produce and assess proposed patches. Those patches were not accepted automatically. Human researchers reviewed each change before it was submitted, checking for regressions and confirming that it addressed the underlying cause rather than only the visible symptom. Even an approved code change does not finish the job. Teams must identify the affected images, rebuild them with the corrected dependency and test the result before deployment. In a poorly documented environment, locating every instance may take longer than producing the patch itself. Accurate inventories give automated tools something concrete to work with. They connect a newly disclosed flaw to the package version, image and workload that actually require attention. Finding the flaw may no longer be the slowest step AI is speeding up code analysis for both attackers and defenders, but many delays still occur after a vulnerability has been identified. One team may spend hours opening images and checking package lists by hand. Another can search a current inventory and see almost immediately which workloads contain the affected version. That difference has little to do with the sophistication of the discovery tool. It comes from decisions made earlier about software inventories, image composition and how containers are built and replaced. As vulnerability research moves faster, the practical advantage belongs to organisations that can establish exposure and deploy a tested repair without first trying to reconstruct what their systems contain. The post How AI is changing the vulnerability response timeline appeared first on AI News. View the full article
  22. Meta is releasing Muse Glimmer under an Apache 2.0 licence for local AI agents that can run on a consumer GPU. The company’s Superintelligence Labs has released the 30-billion-parameter model’s weights on Hugging Face. Meta says developers can use it for local coding, function calling, local agents, and LLM-as-a-judge evaluation. The release targets an operational constraint facing AI teams: cloud-hosted models need network access and central infrastructure. Meta instead pitches Muse Glimmer for workloads that require an on-device model, including personal agents with access to schedules, messages, files, and other private context. Meta Muse Glimmer leads several agent task benchmarks Meta’s benchmark tests put Muse Glimmer ahead of Gemma4-31B and Qwen3.6-27B on five of eight general-agentic benchmarks. The model scored 75.5 on MCP Atlas. Gemma4-31B reached 54.2, and Qwen3.6-27B recorded 62.5. DeepSearch QA follows a similar pattern. Meta reports a score of 74.6 for Muse Glimmer, against 61.7 for Gemma4-31B and 71.1 for Qwen3.6-27B. The supplied announcement identifies both benchmarks as tests of an agent’s ability to work within scaffolds and complete multi-turn requests. The model scored 23.5 on τ²-Banking. Gemma4-31B recorded 15.1. Qwen3.6-27B reached 16.7. Muse Glimmer also posted 47.6 on WildClawBench, ahead of Gemma4-31B’s 37.6 and Qwen3.6-27B’s 43.2. Its GAIA2 result reached 43.3, compared with 36.4 and 40.0 respectively. Other agent scores favour Qwen3.6-27B. Meta’s table gives that model 1,141 on GDPval-AA, against Muse Glimmer’s 953 and Gemma4-31B’s 811. Qwen3.6-27B also led SkillsBench with Skills at 46.6, where Muse Glimmer recorded 44.3. OSWorld-Verified produced the largest gap in this group. Meta reports 75.6 for Qwen3.6-27B. Muse Glimmer reached 65.9, and Gemma4-31B scored 58.5. These tests measure constrained tasks. They do not demonstrate how a local agent will behave after an organisation connects it to its own files, calendars, messaging systems, or internal tools. Coding results split between Muse Glimmer and Qwen Muse Glimmer’s coding results show a narrower comparison. The model led SWE-Bench Pro with a score of 51.2. Meta reports 36.9 for Gemma4-31B and 50.2 for Qwen3.6-27B. SciCode produced a close result. Muse Glimmer scored 43.6, marginally above Gemma4-31B at 43.4. Qwen3.6-27B recorded 39.8. Qwen3.6-27B led two other coding evaluations. It scored 77.2 on SWE-Bench Verified, compared with Muse Glimmer’s 76.0. TerminalBench 2.1 gave Qwen3.6-27B a score of 60.7; Muse Glimmer reached 51.7, and Gemma4-31B posted 43.4. A local coding agent does more than produce code. It needs a scaffold that decides which repositories, terminals, test environments, and commands the model may access. Meta says Muse Glimmer supports OpenClaw and other agent-orchestration patterns, with custom scaffolds covered in its developer documentation. An organisation evaluating the model for software work should define the commands and repositories available to the agent before measuring task success. The supplied material describes retry training for failed tool calls. That behaviour requires controls over repeat attempts, especially where a tool can alter source code or invoke an external system. Multimodal scores favour Qwen in most tests Muse Glimmer accepts interleaved text and images through a dedicated perception encoder. Meta says this design lets agents interpret screenshots, charts, and documents as part of a conversation. The benchmark chart puts Muse Glimmer ahead on Charxiv Reasoning. Its score reached 78.8, against 77.7 for Gemma4-31B and 78.4 for Qwen3.6-27B. Qwen3.6-27B led ScreenSpot Pro with 76.1. Muse Glimmer recorded 75.4, and Gemma4-31B scored 75.9. The same model led OmniDocBench v1.5 at 77.8, compared with Muse Glimmer’s 75.8 and Gemma4-31B’s 72.5. MMMU Pro produced smaller differences. Meta lists Muse Glimmer at 74. Qwen3.6-27B reached 75, and Gemma4-31B posted 73. These results matter for teams considering agents that act on visual interfaces. A screenshot-reading model can interpret what it sees, yet local testing must still cover permissions, display layouts, document formats, and errors returned by connected tools. Safety figures show lower reported attack success than Qwen Meta also reports two safety-related evaluations: CI Memories and Siren AgentDojo. The chart uses different measures for each test. On CI Memories, Meta lists a violation rate of 26.4 for Muse Glimmer and a coverage score of 64.8. Gemma4-31B recorded a violation rate of 12.1 with coverage of 53.0. Qwen3.6-27B posted a violation rate of 53.4 and coverage of 66.9. The Siren AgentDojo result uses attack success rate and utility. Meta gives Muse Glimmer an attack success rate of 28.4 and a utility score of 94.2. Gemma4-31B scored 25.6 on attack success rate, with utility at 90.8. Qwen3.6-27B recorded 40.3 and 92.7. General reasoning results add context to agent claims Muse Glimmer led four of six general-capabilities-and-reasoning tests in Meta’s comparison. It scored 77.0 on IFBench. Gemma4-31B recorded 76.0, and Qwen3.6-27B reached 70.8. The AIME 2026 score was 94.7 for Muse Glimmer. Meta reports 89.2 for Gemma4-31B and 94.1 for Qwen3.6-27B. On AA-LCR, Muse Glimmer reached 80.0, ahead of 68.3 and 73.3. The model also led Beam 128K at 65.1. Qwen3.6-27B scored 63.0. Gemma4-31B recorded 58.2. Gemma4-31B led GPQA Diamond with 85.7. Muse Glimmer scored 83.5, followed by Qwen3.6-27B at 84.2. Gemma4-31B also took the top score on Humanity’s Last Exam, Text No Tools, at 23.6; Muse Glimmer reached 22.0. One model does not lead every test. Meta’s results instead show Muse Glimmer competing closely with two similarly sized models across a mixed set of agent, coding, visual, safety, and reasoning evaluations. Memory limits shape the local deployment design Meta says a full-precision 30-billion-parameter model would require more than 55 GB of memory. Muse Glimmer instead uses approximately 4-bit weight quantisation, reducing the language model to under 20 GB. That allocation leaves memory for a KV cache. The model also needs room for its perception encoder and a speculative-decoding drafter. Meta targets a 24 GB or 32 GB memory envelope for these components. The company says the DFlash-based drafter proposes blocks of tokens for the main model to verify in parallel. Meta says this speeds generation compared with standard token-by-token output and retains identical output quality. The supplied post does not include token-per-second figures, prompt sizes, power data, or concurrency results. Meta tested its K-Quant-17GB version with the quantised DFlash drafter on MacBook M4-Max hardware, MacBook M5-Max hardware, and an RTX-5090. It describes the resulting experience as suitable for fluid conversation and real-time agent interaction. The public weights are available through Hugging Face. Meta says integrations with llama.cpp, MLX, and ExecuTorch will arrive in the coming days. See also: Alibaba tests new business model for Qwen open-source AI Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post Meta Muse Glimmer brings local AI agents to consumer GPUs appeared first on AI News. View the full article
  23. Physics AI can now explore thousands of design variations in the time it would take a traditional simulation to chew through a handful of them. Precisely up to 1,000 times faster, according to Siemens. What it cannot do is sign off a safety-critical part. On that, the technology has a firm limit, and Sam Mahalingam, who leads the business building it at Siemens Digital Industries Software, states it without hedging. “Is this good for safety-critical applications?” he said, on the sidelines of Realize LIVE Asia-Pacific in Bengaluru. “No, it is not.” That matters because the answer cuts against two years of an industry insisting its AI can do nearly everything. The value for engineers is not in the speed Siemens is selling, but in knowing exactly where that speed stops being safe to rely on. What physics AI actually does, and what it does not The technology in question is Simcenter PhysicsAI, Siemens’ geometric deep-learning software, which the company says can make design predictions up to 1,000 times faster than a traditional solver. The mechanism matters to understanding the caveat. Rather than computing the physics from scratch each time, a surrogate model learns from historical simulation data and predicts the outcome for a new design. It is an estimate, produced in seconds, not a full calculation. The obvious worry is accuracy, and Mahalingam met it head-on. For decades, he explained, engineers benchmarked physics-based simulation against physical testing until the correlation was tight enough to trust. AI is now being measured against that same physics baseline. “What we are seeing is that if you have sufficient data, it is very close to a physics-based solver,” he said, the 1% to 3% variation Siemens cites in its own case studies. Close, but not close enough to certify a life-or-death component. And this is where Mahalingam departs from the standard vendor script. The surrogate is not a replacement for validation; it is a filter placed in front of it. “You explore a lot more design variations using this faster engine, the physics AI surrogate model, zero in on two or three designs that you feel are good, that you can further do detailed design on using a physics-based simulation,” he said. Only once those finalists clear a full physics-based check does a design move toward manufacturing. He was explicit that even a marquee example, a Continental airbag case Siemens has showcased, sits inside that boundary. “This is for the initial design exploration,” he said. “It is not that you are only validating with physics AI and you are saying, okay, I’m going to go recommend that design for manufacturing. No, that’s not the case.” The dependency the speed numbers do not mention There is a second limit that the acceleration figures tend to obscure, and it surfaced when the conversation turned to how these models are trained. Several of Siemens’ headline results, including cases involving Magna and Continental, rest on AI trained on synthetic data: simulation output generated by Siemens’ own solvers rather than real-world measurement. If the AI only learns from the simulation, the question is whether it can ever be better than the simulation that taught it. Mahalingam did not dodge the circularity when asked. In Magna’s case, he said, the customer ran a broad design exploration in Simcenter HEEDS, Siemens’ design-search tool, and solved the variations at speed using Simsolid, a solver that skips the slow mesh-building step. That simulation output was then fed back into training the physics AI model. Where a customer has no data to begin with, “they first generated synthetic data with Simsolid and HEEDS, and then they went back, took that data, trained a physics AI model.” The surrogate, in other words, is only ever as good as the simulation beneath it, a constraint he acknowledged rather than waved away. What keeps that from becoming a trap, he argued, is a guardrail built to stop the model predicting on ground it has never seen. A surrogate trained on variations of one shape will fail if asked to predict a radically different one, and it is designed to say so. “We have put in guardrails where it comes back and says, hey, I cannot predict this. This is completely a different shape compared to what you trained it on,” Mahalingam said. “So the engineer cannot shoot themselves in their own legs.” Why the honesty is the story The candour is not self-effacement; it is positioning. Every simulation vendor is now racing to attach AI to its portfolio, and the credibility risk is that buyers stop believing any of the numbers. By marking the edge of the technology–safe for exploration, not for final sign-off; powerful with data, useless beyond its training envelope–Siemens is making a bet that engineers trust a tool more when it tells them what it cannot do. It lands differently coming from the simulation side of the house. Chip-design and enterprise-AI vendors have spent the hype cycle promising autonomy; a company whose customers model ****** structures and jet engines is instead insisting that the human validation step stays exactly where it is. That is not a hedge against AI. It is a clearer-eyed account of where it belongs, as the fast first pass that widens the search, with the physics-based solver still holding the pen on anything that has to be right. That is a narrower claim than the market is used to hearing, and a more durable one. Siemens is selling the 1,000x, but the more valuable thing it is offering engineers is the boundary around it. See also: Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post The limits of physics AI: where Siemens says the human stays in charge appeared first on AI News. View the full article
  24. Stanford researchers have synthesised nearly 300 phages from DNA sequences produced by the Evo 2 generative AI model. Laboratory testing narrowed the group to 16 phages that showed particularly strong E. coli-killing activity. The work centres on bacteriophage ΦX174, pronounced “FYE-ex-1-7-4”. Brian Hie, an assistant professor of chemical engineering and Dieter Schwarz Foundation Stanford Data Science Faculty Fellow, created Evo 2 with bioengineering graduate student Samuel King leading the experimental work described in the paper. Evo 2 takes phage genomes into the laboratory Evo 2 generates new DNA sequences from a small starting snippet of a phage genome. The researchers asked the model to produce an entire ΦX174 genome in one left-to-right pass. “In this case, we wanted the model to generate the entire genome end-to-end in a single left-to-right pass. We didn’t add anything,” Hie said. The process generated thousands of candidate genomes before the team selected sequences for chemical synthesis and laboratory testing. ΦX174 offered a relatively compact test system. Its genome contains fewer than 6,000 base pairs, compared with roughly 3 billion base pairs in the human genome. Hie said researchers still face a difficult task when interpreting even a 5,400-character DNA sequence gene-by-gene. Hie said some of Evo 2’s suggested phages showed higher fitness than native ΦX174 in laboratory testing. The project therefore tests whether a model can create entire viable viral genomes, rather than only proposing local DNA edits. Candidate screening comes before DNA synthesis King developed a computational framework to reduce the number of candidate genomes sent for synthesis. The framework assessed traits drawn from ΦX174 and related phages before the team selected options for laboratory work. DNA synthesis sets a practical constraint. The researchers generated genomes with Evo 2, evaluated them against their design criteria, then chemically-synthesised selected candidates and tested which genomes performed best in the lab. “One of the main parts of the design framework was figuring out what traits the genomes should have based on ΦX174 and related phages,” King said. “The framework involved several key steps: generating genomes using Evo 2, evaluating options based on the design criteria, selecting optimal candidates, synthesising them chemically, and then testing them in the lab to see which genomes worked best.” Hie said the framework reduced synthesis costs by concentrating spending on the candidates his team judged most viable. This sequence also defines the operational boundary of the result. Evo 2 generated thousands of possibilities, yet the researchers still required computational evaluation, chemical synthesis, and laboratory assays to identify viable phages. Resistance testing centres on a 16-phage mixture The researchers selected more than one E. coli-targeting phage because bacteria can develop resistance to a single treatment. Hie said phage mixtures could make it harder for bacteria to evade every member of a treatment. “If the bacteria gains resistance to a single phage, it’s game over for the medication,” Hie said. “But if you have multiple genetically distinct phages in a mixture, it would be harder for the bacteria to develop resistance to the entire *********.” Stanford reports that a ********* containing the 16 selected phages rapidly overcame resistance in E. coli that was immune to native ΦX174. Hie said similar work could pursue phages aimed at methicillin-resistant Staphylococcus aureus, or MRSA. He also named Pseudomonas aeruginosa, which Stanford describes as a leading cause of medically-resistant infections acquired in hospitals. Open-source access extends the research programme Hie has released Evo 2 as open-source software. Researchers can download the model and use it to design genomes. The release has raised safety and security discussions, according to Stanford’s account. Hie acknowledged that bad actors could modify versions of the tool, though he argued that existing pathogens create a greater risk because people can access and produce them more easily. Hie also said AI-enabled systems can support responses to naturally-occurring pandemics and provide defence options against man-made biological threats. Those views reflect his assessment of the tool’s potential uses and risks. King described the research benefit in narrower terms: “One of the most rewarding parts of this project is the creativity Evo 2 allows. New doors in science are now open because of what we can do with these models.” The next phase will extend Evo 2 to longer and more complex DNA, according to Stanford. Hie is working with researchers at Stanford and elsewhere on additional bacteriophage designs. Small bacterial genomes could also become a target for the model. Stanford says those genomes might support engineered microbes designed to produce chemicals, medicines, or fuels. Hie framed the remaining technical work around two questions: “The biggest open questions for me are how do we get greater genetic novelty and how do we get greater controllability of the outcomes?” See also: Why health AI interfaces must adapt to user expertise Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information. AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here. The post Stanford Evo 2 AI model generates phages against E. coli appeared first on AI News. View the full article
  25. Every post you see, every Reel that autoplays, and every ‘Explore’ page suggestion on Instagram is now decided by its AI system. With over three billion people using it, that’s not a small detail; it’s the whole algorithm. But here’s the point: the more AI controls over how you see the content, the more it rewards what feels human. As said by Instagram head Adam Mosseri in a year-end memo, that platform will prioritise “real human-made content” over AI-generated content. So, is AI replacing the human element of Instagram, or protecting it? The honest answer is neither extreme. AI has taken over the mechanics of engagement. Humans still own the meaning behind it. Here’s how that split actually works. What is the role of AI in Instagram engagement? The simple answer is AI decides what gets shown. Yet, real human interaction decides what grows on the platform. In fact, the Instagram engagement rate in 2026 sits at just 0.7%. making it important to understand each engagement more than ever before. Instagram doesn’t operate on one algorithm. Instead, each part of the app has its own, like one for Feed, one for Reels, one for Stories, and one for the Explore page. Each one predicts specifically on the basis of what you enjoy based on your past behaviour. That’s why two people can open Instagram at the same moment and see completely different content. Instagram signals that control the reach and engagement Three signals control the reach and engagement: watch time, likes per reach, and sends per reach. When someone shares your post to a friend, it carries trust points. An algorithm can guess what you’ll like, but it’s impossible to fake someone caring enough to send it along. That’s where human touch helps the algorithm most. Initial interaction with the post and Reel also matters. A Reel that gets good reach in its initial 30 minutes will often outrank one that was posted three days ago and is growing gradually. In algorithmic language, that’s called engagement velocity, and it helps content that gets viral with immediate and genuine reaction. With AI evolution, Instagram is getting better at spotting what is fake. Fake followers, bot comments, and coordinated engagement pods get flagged and penalised. The same algorithm that personalises your feed is also a tool against fake engagement and reach. What can automation manage behind the scenes on Instagram? With AI integration into the system, now people can spend more time on strategy, growth, and creativity instead of repetitive production work. There was a time when growing on Instagram was only about writing captions, resizing images, planning video variations, and scheduling reels at the best times. All of this was eating up important hours every week. AI Instagram automation tools can now do the same in just a few minutes. For instance, brands use AI to draft first versions of captions, generate on-brand visuals, repurpose one video into five formats, and even predict which audience segment will respond best to a specific post. That’s where creators remove the busywork hurdle and can focus on parts that actually require a human touch: strategy, voice, story, and judgement calls. All this helps, but there’s a limit to this. Instagram now suspends accounts that repost content too often. Automation can help with speedy production, but it can’t replace what a real human brain thinks, and Instagram’s AI system is designed to notice the difference. How is AI transforming customer communication in the Instagram DMs? AI chatbots now handle the first reply, but complicated or sensitive conversations still get passed to a real person. Direct messages have quietly become one of the busiest channels on Instagram. People don’t email a business anymore. Instead, they DM it. And they expect a fast answer. According to a study by Harvard Business School, AI chatbots helped human agents respond to customer DMs some 20 percent faster. That’s the difference that helps businesses engage with customers in a more efficient way. Two types of chatbots exist: Rule-based bots that trigger a set reply when someone types a specific word like “order status” or “size chart” AI agents that actually understand customer demand, hold a natural conversation according to the queries asked, and answer even unusual questions using the brand’s own information. With time, more advanced AI agents can handle almost anything customers ask for. They help you with order status, issue a refund, solve a dispute, all using live data. That used to depend on a support team with top-notch communication skills. But now it happens anytime, in any time zone. This AI integration is helping Instagram turn into a customer favourite marketplace, not just a social media app. But almost every serious version of these chatbots includes one key feature: human handoff. When a conversation gets complicated, emotional, or high-stakes, the AI steps back, and a real person takes over. That handoff isn’t a weakness in the system, it’s the protection of human rights. One more important fact is that no business should be using AI as a one-size-fits-all solution. What’s coming ahead? Expect AI to keep expanding its reach into production and support, while Instagram keeps tightening the rules around what counts as genuine. AI integration into production and support systems is obvious. At the same time, Instagram’s algorithm is focusing on what counts as real. A few patterns are already visible: Not just DM automation. Initially, AI agents were designed to answer questions, but real customer service resolves the questions. So, now better AI systems are helping businesses connect directly to what matters for them. More focus on genuine content. Instagram’s focus on “real” content isn’t a one-time announcement. Expect more filters, more detection systems, and more penalties for accounts that lean too hard on AI. Who manages it well will be the new skill. As more brands use AI for captions, replies, and content, the real difference won’t be who AI has. It’ll be who manages it well. In short, we can say that tools are becoming better day-by-day. Hence, the brands which will win are the ones who use them. Instead, those will be ones who use it well and alter the tone according to their brand voice and values. Final words On the bottom line, brands that understand the system are using one to protect the other. AI didn’t remove the human touch from Instagram; in fact, it helps humans handle the job in a more efficient way. AI now runs the ranking systems, produces multiple drafts, and answers the basic questions, while humans show up in the moments that actually build trust. That combination: AI for scale and humans for meaning is exactly what Instagram’s algorithm is designed to reward in 2026. The post How AI Is changing Instagram engagement without replacing the human touch appeared first on AI News. View the full article

Important Information

Privacy Notice: We utilize cookies to optimize your browsing experience and analyze website traffic. By consenting, you acknowledge and agree to our Cookie Policy, ensuring your privacy preferences are respected.