AI term
On-premises inference
What is On-premises inference?
Running a trained model to generate outputs on computing hardware located at the organization’s own facilities.
In other languages
- 한국어사내 추론
- 조직 내부의 컴퓨터 장비에서 학습된 모델을 실행해 결과를 생성하는 과정입니다.
- 日本語社内設置型推論
- 組織内に設置した計算設備で学習済みモデルを実行し、予測や生成を行う処理です。
Related Terms
- Local InferenceThe process of running AI models directly on a user's own hardware rather than relying on remote cloud servers.
- Inference computeComputational resources used to generate outputs from a trained machine-learning model.
- Inference HarnessThe software infrastructure and environment responsible for loading a pre-trained machine learning model and managing its execution to produce outputs.
- Inference ScalingTechniques that improve performance by increasing the computational budget during the model's inference phase.
- Test-time scalingA method of improving model outputs by increasing computational resources or processing time during the actual usage or inference phase.
- Inference efficiencyMeasure of a model's computational performance during the prediction phase, typically involving speed and resource usage
- Sim2RealThe process of transferring machine learning models trained in a simulated environment to operate successfully on physical hardware in the real world.
- Federated learningA machine learning technique that trains models across multiple decentralized devices holding local data without exchanging them.