Enterprise generative AI has hit a familiar ceiling. The initial phase of prototyping with off-the-shelf, proprietary API endpoints delivered fast proof-of-concepts, but moving multi-agent systems into production has exposed serious friction points.
IT leaders are running headfirst into volatile token costs, rigid data boundaries, and black-box inference systems that fail strict compliance reviews.
Deloitte’s launch of its global Open Model Engineering practice addresses this exact pivot.
By helping enterprise and government clients build, fine-tune, and run open-source AI frameworks alongside closed-source tools, the consultancy is validating a major market shift: enterprise AI maturity requires an infrastructure that companies can actually control and audit.
Modernizing the AI Stack Through Custom Infrastructure and Token Control
For enterprise workloads, relying entirely on frontier proprietary models is often an over-engineered and costly choice. Routing every standard business task through massive closed models inflates API expenses and introduces latency risks.
Deloitte’s practice focuses on building mixed-model architectures where open models handle high-volume, specialized tasks, while proprietary platforms remain available for edge cases requiring brute-force reasoning.
The technical backbone of this rollout leans heavily into NVIDIA’s ecosystem. The practice focuses initially on the NVIDIA Nemotron family of open models and NVIDIA NIM microservices, integrating them with platforms like the Zora AI digital workforce engine.
This setup gives engineering teams the tooling to deploy multi-agent workflows inside their own private environments rather than relying on external hosted endpoints.
Crucially, Deloitte is pairing this infrastructure with dedicated personnel. By hiring, training, and deploying certified forward-deployed engineers directly inside client environments through fiscal year 2027, the firm is acknowledging that running open models is fundamentally a systems-engineering discipline.
It requires internal pipelines for quantization, model serving, prompt routing, and continuous evaluation capabilities that most internal IT departments cannot build overnight.
Securing Data Sovereignty and Operational Transparency
Beyond raw infrastructure cost, the primary driver pushing organizations toward open model engineering is regulatory necessity and intellectual property protection.
When companies operate in healthcare, financial services, or critical public infrastructure, feeding proprietary data and proprietary trade secrets into external commercial model clouds introduces compliance liabilities that legal departments are no longer willing to sign off on.
Open models resolve the sovereignty dilemma by enabling true localized hosting. Organizations can run enterprise agents either on-premises or within isolated sovereign cloud boundaries, ensuring raw customer data, proprietary training corpuses, and generated outputs never cross jurisdictional or organizational borders.
This allows teams to fine-tune models to match regional languages, cultural context, and industry-specific terminology without forfeiting custody of their training weights.
Furthermore, running open architectures introduces transparency at the inference layer. When an AI agent makes a decision or accesses a database, engineers have direct visibility into model behaviors, weights, and agentic harnesses.
Being able to inspect, patch, and harden an AI model against hallucinations or security exploits gives technical teams an audit trail that closed APIs cannot match.
The move away from a single-vendor AI pipeline toward a modular, mixed-model environment is not just an infrastructure trend; it is the baseline for sustainable enterprise deployment.
Real competitive advantage will not come from calling the same commercial API as everyone else, but from owning the pipeline, fine-tuning domain-specific models, and governing the operational data from the ground up.
Source: Official Deloitte, "Deloitte Launches Open Model Engineering Practice"




