What if your business had its own AI?

people in an office working on coding on a computer screen
Dolo Miah

Dolo Miah

CEO & CTO

Published:

For the last few years, the conversation around business AI has mostly focused on what the technology can do. Can it write this, analyse that, automate a process or help a team work faster? Those questions still matter, but as AI moves from experimentation into everyday business operations, I think we're starting to see businesses asking a different set of questions.


How will token pricing impact the commercial viability? What happens to the data we're giving it? How comfortable are we putting commercially sensitive information through somebody else's model? What if models change, could that negatively impact my current processes? And as AI becomes a fundamental technology for our business, how dependent do we want to be on a handful of external providers?


One answer to some of those questions is surprisingly simple: run the AI yourself.


That might sound like something reserved for huge technology companies with enormous computing infrastructure, but it isn't. It is now possible to run scaled-down but capable language models (LLMs) on infrastructure controlled by the business, whether that's its own servers, private cloud environment, edge hardware or, in some cases, an individual’s PC or laptop.


I don't think many businesses have fully appreciated what that could mean yet.


We need to start thinking about the cost of AI differently


The cost of AI isn't particularly noticeable when you're experimenting with it. A few ChatGPT licences or occasional API calls are unlikely to cause much concern, which makes it very easy to prove an idea and start integrating AI into different areas of a business.


The economics start to look different once those experiments become operational systems. If an AI model is processing customer enquiries, analysing documents, supporting employees, monitoring equipment or powering an automated workflow, it could eventually be called thousands or millions of times. Most commercial AI APIs charge according to usage, typically based on the volume of tokens being processed and generated, so greater adoption can mean greater ongoing cost.


Now this ought to have been modelled as part of the business case however what I am seeing is that often the long-term cost of tokens isn’t well considered. But even when it is, there’s still a big unknown in terms of whether the cost of a token can be predicted. When we consider the macro-economics of the AI technology industry we see huge investment to build the datacentres that host and run the models, but also that the AI providers are not yet profitable. That means current token pricing may have to change whether explicitly or by stealth. AI Token Spend Is Rising, But Measuring Value Remains Challenging 


Hosting a model yourself doesn't suddenly make AI free, of course. There is hardware to consider, along with electricity, maintenance, monitoring and the expertise needed to build and operate the system properly. What changes is the nature of the cost. Instead of paying an external provider every time your system uses intelligence, you're investing in infrastructure that you control with a predictable cost forecast.


Whether that makes financial sense depends entirely on the application and scale, but it's an equation I think more businesses should at least be considering as their use of AI grows.


As we dive headlong into AI, the data sensitivity question is ringing louder


Cost is only part of this conversation. For me, data considerations are potentially much more significant.


AI becomes considerably more useful when it understands the context of your organisation. A general-purpose model can help you write an email, but an AI system that can work with your processes, documents, customer information, operational data and internal knowledge has the potential to do much more valuable work.


The problem is that this is also some of the information businesses are most protective of.


This doesn't mean that using a cloud AI provider is inherently insecure. Major providers have invested heavily in enterprise security, privacy controls and agreements around how business data is handled. For many organisations and many applications, those protections will be perfectly appropriate.


But there is a difference between trusting another organisation to handle your data securely and designing a system where that data doesn't need to leave your environment in the first place.


If you're dealing with sensitive intellectual property, regulated information or commercially valuable data, that distinction can become important. Local models provide an additional control plane enabling the use of AI and agents whilst fully assuring customers on confidentiality considerations. 


For example, a company we’ve recently been helping consider their options wants to enable AI to help accelerate their sales and quoting process which uses commercially confidential specifications provided by their client and internal reference data such as labour and materials costs.


Edge AI - when AI in the cloud isn’t the right fit


This is where I think the conversation becomes particularly interesting as new architecture options become possible.


The world today is increasingly digitised which means that data is everywhere. Collecting and sending that data to the cloud for intelligence to be applied to it comes with network, latency and cost hurdles. With a local language model, you can bring the intelligence closer to where the data is generated and used, minimising round-trip delays and the cost of network bandwidth.


Running AI locally means you can put intelligence in places where relying on a constant connection to a cloud model isn't practical - this is often the case for businesses operating in highly distributed or harsh environments.


Consider a manufacturing environment, a remote piece of infrastructure such as a pump or wind turbine, a ship, an engineering site or any number of situations where connectivity is unreliable, latency matters or moving large volumes of data backwards and forwards simply doesn't make sense.


Having a localised AI means a model can collect, analyse and act on data where it is generated and used. This enables new levels of automation to identify a significant event (e.g. a failing pump) and act on that event in real-time without reliance on a cloud AI service.


This is the key ability of Edge Computing and we’ve talked about it in this way for years. What’s changing the game is that small and specialised language models can be optimised to run on modest edge devices, bringing a new era of intelligent capability to enterprises that need to manage their distributed and physical assets ever more closely.


There is also a question of model control


Businesses are currently building a lot of AI functionality on infrastructure they don't own and models they don't control. There's nothing inherently wrong with that; using established platforms is often the quickest and most sensible way to get something working.


But the more important AI becomes to an organisation, the more important those dependencies become too. Your application might rely on a particular provider's pricing, availability, model behaviour, usage policies and future product decisions. You will need to be able to control the impact of providers changing any of those things.


Using open models on your infrastructure doesn't remove every dependency or technical challenge, but it can mitigate many of the risks. You can make decisions about where the model operates, what information it can access, how it connects to other systems and, depending on how you've designed the solution, whether you want to change models in the future.


I suspect that control will become increasingly critical as businesses move beyond isolated AI experiments and AI becomes a part of their foundational technology infrastructure.


So, should every business have its own AI?


Probably not, and I wouldn't recommend self-hosting an AI model simply because you can.


There are plenty of situations where a commercial cloud model will be cheaper, easier and more capable. Running AI yourself introduces different considerations around hardware, maintenance, model performance, security and technical expertise, and the largest models still require substantial computing resources.


There are also increasingly interesting hybrid approaches. A business might use a local model for sensitive or repetitive tasks while using a larger cloud model when it needs capabilities that aren't practical to run locally. The decision doesn't have to be entirely cloud or entirely private.


That's why I don't think the interesting question is really “Should we have our own AI?” The better question is “Where should our AI run?”


The answer should come from the problem you're trying to solve. What information does the system need access to and where does it reside? How sensitive is the data? How frequently will the model be used? Does it need to operate without connectivity? What level of capability and performance does the task require? What will the system cost to operate at the scale you're aiming for?


Once you consider those things, you can make a much more informed decision about the architecture of your solution. For some businesses, the answer will still be the cloud. For others, running some AI workloads within their own infrastructure could solve very real concerns around cost, confidentiality, connectivity, performance and control.


This is no longer a theoretical option reserved for organisations with enormous computing resources. Many open models have smaller versions which have had techniques such as quantisation applied with reduced numbers of parameters which means they can run on modest hardware such as a laptop or edge device. Why we still need small language models – even in the age of frontier AI.


With considered selection of the right small model and appropriate tuning, acceptable results are absolutely possible using local infrastructure.


Thinking about where AI could fit into your business?


Whether you're considering cloud, local or hybrid AI, the right approach depends on what you're trying to achieve, the data involved and how the technology needs to operate in practice. We can help you work through the options and build an AI solution around the specific needs of your business.


Talk to us about your AI project →

Dolo Miah

Dolo Miah

CEO & CTO

Dolo is the CEO & CTO of New Icon, with more than 30 years’ experience across technology, enterprise architecture and digital transformation. He works with organisations to unlock greater value from their data, systems and emerging technologies, with particular expertise in AI, edge computing and real-time digital transformation.

Reimagine your digital future today

Send us a message for more information about how we can help you and your business

Reimagine your digital future today

Send us a message for more information about how we can help you and your business

Reimagine your digital future today

Send us a message for more information about how we can help you and your business

Reimagine your digital future today

Send us a message for more information about how we can help you and your business

Services

Capabilities

About

Linebreak

New Icon is a Linebreak company

© Newicon Ltd. Registered in England and Wales. Company No: 05904359 | VAT: GB 993768447.

Designed and built by New Icon in Bristol, a Linebreak company.

Linebreak

New Icon is a Linebreak company

© Newicon Ltd. Registered in England and Wales. Company No: 05904359 | VAT: GB 993768447.

Designed and built by New Icon in Bristol, a Linebreak company.

Linebreak

New Icon is a Linebreak company

© Newicon Ltd. Registered in England and Wales. Company No: 05904359 | VAT: GB 993768447.

Designed and built by New Icon in Bristol, a Linebreak company.