How do you decide what model to use? Let’s look at workload, deployment, regulated environments and other criteria for selecting the right AI model.
In an AI application, a model is the trained component that receives an input and produces generated output such as text or structured data. The models we’re considering are large language models (LLMs), which can interpret language and return text or structured data (e.g., GPT, Claude, Llama, Mistral).
A model can perform well in a prototype and still be the wrong choice for production because the surrounding requirements change what counts as a good fit.
For example, a security review may show that requests cannot cross a regional boundary, while load testing may reveal that the model misses the feature’s response-time budget. Cost at projected traffic and the team’s ability to operate the deployment can also remove models from consideration.
In this article, we’ll look at the practical criteria that shape model selection. We’ll focus on workload fit and deployment requirements, including the constraints that come with regulated environments.
What the Model Needs to Do
Before we compare models, we need a clear description of the job. A strong score on a broad reasoning test won’t tell us whether a model can handle our document formats or return tool arguments that our backend accepts.
An evaluation set built from real inputs gives us a consistent way to compare candidates. Each test needs an expected result or a scoring rule so we can judge every model against the same requirements.
For each model, we can ask four questions:
- Does it produce useful and accurate output?
- Can it handle the required context length and languages?
- Do its structured responses match the schemas our application expects?
- Do its tool calls select the right operation and provide valid arguments?
A smaller model may be enough for document classification, while an agent choosing tools may need stronger reasoning. Public benchmarks and model cards can help us find candidates before we run these tests. A model card describes intended uses and limitations along with licensing and published evaluations. Those results give us a starting point while our own inputs show whether the model fits the feature we intend to ship.
Where the Model Runs
Once we know a model can do the job, the next question is where it will run. Running a model to produce an output is called inference, and the infrastructure that handles this work becomes part of the selection decision.
With a hosted model API, the provider operates the serving stack. A self-hosted model makes our team responsible for that stack in a cloud account or local environment.
A self-hosted model can run in our cloud account or our own data center. When it runs in our data center, the deployment is on-premises.
There are four common arrangements, each giving us a different level of control and operational responsibility:
If we choose to self-host, we need access to the model’s weights (the numerical values learned during training that help determine its output) under a license that allows our intended use. We also need to confirm that the model works with serving software our team can maintain and fits within the GPU capacity available at normal and peak traffic.
Because deployment affects performance and cost, we should compare candidates in the setup we intend to use. A hosted endpoint and a self-hosted deployment of the same model can differ in response time and total cost at production traffic.
Model Selection in Regulated Environments
The deployment choices become more specific when the feature handles regulated data. The applicable rules and our organization’s risk assessment give us concrete requirements for where inference can run and how its data must be handled.
In healthcare, HHS guidance on HIPAA and cloud computing says a covered entity may use a cloud service to process electronic protected health information when the required business associate agreement and HIPAA safeguards are in place. The organization still needs to understand the cloud environment and complete its own risk analysis.
For broker-dealers, FINRA’s cloud guidance says moving infrastructure to the cloud does not remove the firm’s regulatory responsibilities. The deployment still has to support vendor oversight and recordkeeping.
That means we need more than a model that performs well. Before using one with regulated data, we should be able to answer some additional practical questions:
- Where are prompts and outputs processed, and can the provider use them for training?
- How is the data retained and deleted, and which access and audit records can we export?
- Which contracts govern the provider and any subprocessors it uses?
- How are model changes approved, and can we restore an earlier version?
Making the Final Choice
Once we have a shortlist that fits the workload and deployment requirements, we can make the final choice by answering four questions:
- Does the model meet the workload? The model needs to clear the quality threshold on representative inputs and support the capabilities the feature uses.
- Can we deploy it within the required boundary? Its license and contracts must permit our intended use, and its data handling has to satisfy the relevant policies.
- Can it meet the runtime budget? Response time and total cost need to remain within budget at normal traffic and peak demand.
- Can we support the deployment? The responsibilities should be clear, including who manages capacity and model changes.
Together, these questions keep us from choosing a model based on quality alone. Recording the answers also shows why it was selected and gives us a useful starting point when the workload or deployment environment changes.
Wrap-up
Choosing a model begins with the work it must perform and where inference can run. From there, we can compare the remaining candidates against the response-time budget and the cost of the intended deployment.
A hosted model can still be a good fit for a regulated workload when its contracts and controls satisfy the requirements. When inference needs to stay within infrastructure we operate, self-hosting can provide the necessary control. The right model is the one that fits both the feature and the environment in which we need to run it.
For more on building AI-powered applications and agents with Progress, check out the following resources:
- Agent Engineering Platform | Progress Telerik
- Progress Forge - AI-Assisted Software Development
- Progress Agentic RAG











