Inference: meaning and usage
Official source
What is Inference
Inference is the process in which a trained model receives an input and produces a prediction, generated response, or other output. It is distinct from training the model's parameters.
Detailed explanation
What it is, what it does, and where it fits
Inferenceは、学習済みの機械学習モデルへ新しい入力を与え、予測や生成の結果を計算する処理です。Googleの機械学習用語集は、LLMでは入力プロンプトに対して応答を生成すること、一般の機械学習ではラベルのない例へ学習済みモデルを適用して予測することとして説明しています。ユーザーの要求ごとに結果を返すOnline InferenceやDynamic Inferenceと、複数の入力をまとめて処理して結果を保存するBatch Inferenceがあります。Inferenceの速度や費用は、モデルの大きさ、入力と出力のトークン数、ハードウェア、バッチサイズ、量子化、キャッシュ、ネットワークなどに左右されます。生成結果が返ったことは、内容が正しいことや、業務処理が完了したことを意味しません。モデルの学習、Fine-tuning、評価、Servingは関係しますが、Inferenceはその中の実行・予測段階として区別して扱います。
本番システムでは応答時間、スループット、メモリ、費用、再現性を測る必要があり、学習とは別の設計課題になります。
現在はバッチ処理、ストリーミング、量子化、キャッシュ、アクセラレーターを組み合わせてサービス要件に合わせます。
Related terms
Terms that help place it in context
Sources and verification
Verification date and sources
Official source
Last verifiedAug 23, 2026
This entry is based on researched sources and item-level verification notes.
Understand the term,
then check
the context.
Expansions, meanings, domains, and evidence are shown separately so an abbreviation can be read in context.