AI English Shift · Lessons / F15

Choosing and operating an AI system

选择与运行 AI 系统

Open the interactive lesson · 打开互动课程

Model selection starts with the task and its constraints, because a model's general reputation does not establish its suitability for a particular workflow.

选择模型应从任务及其约束出发,因为一个模型的普遍口碑并不能证明它适合某个具体工作流程。

モデル選定は作業と制約から始まります。一般的な評判だけでは、特定の業務への適性は分からないためです。

Useful constraints include acceptable errors, response time, supported languages, data handling, available tools, and operating cost.

值得考虑的约束包括可接受的错误、响应时间、支持的语言、数据处理方式、可用工具和运行成本。

有用な制約には、許容する誤り、応答時間、対応言語、データの扱い、利用できるツール、運用費用があります。

A hosted model service and a self-hosted model place different responsibilities on the team, including deployment, access controls, scaling, and maintenance.

使用托管模型服务与自行托管模型,会让团队承担不同的责任,包括部署、访问控制、扩缩容和维护。

ホストされたモデルサービスと自社運用のモデルでは、配備、アクセス制御、拡張、保守など、チームが担う責任が異なります。

The label open weights describes access to model weights, but it does not alone establish every permission or operating requirement.

“开放权重”描述的是可以获取模型权重,但仅凭这个标签,并不能确定所有使用许可或运行要求。

オープンウェイトという呼び方は重みへのアクセスを表しますが、それだけですべての利用許可や運用要件が分かるわけではありません。

Compare candidates on representative tasks using the same evidence, success criteria, and accounting boundaries.

比较候选系统时,应使用有代表性的任务,并采用相同的证据、成功标准和成本核算范围。

候補は、同じ根拠、成功基準、費用計算の範囲を使い、代表的な作業で比較します。

In a fictional pilot, System A spends 12 dollars across 100 attempted tasks and successfully completes 60 of them.

在一个虚构试点中,系统 A 尝试了 100 项任务,共花费 12 美元,其中成功完成了 60 项。

架空の試行で、システムAは100件の試行全体に12ドルを使い、60件を成功させます。

System B spends 18 dollars and completes 90 tasks successfully, so both systems cost 0.20 dollars per successful task before other expenses.

系统 B 花费 18 美元,成功完成了 90 项任务,因此在未计入其他支出的情况下,两者每成功完成一项任务的成本都是 0.20 美元。

システムBは18ドルを使い90件を成功させるため、他の費用を除くと、どちらも成功1件あたり0.20ドルです。

The cheaper total bill does not make A the better system, especially if its failures create additional review or recovery work.

总账单更便宜并不意味着 A 是更好的系统,尤其是当它的失败会带来额外的复核或恢复工作时。

総請求額が低いだけではAが優れるとはいえず、失敗で確認や復旧の作業が増えるなら、なおさらです。

Latency also needs a precise definition, such as time to first visible output or time until the entire usable result is ready.

延迟也需要精确定义,例如从请求开始到首次出现可见输出的时间,或直到完整可用结果准备就绪的时间。

遅延時間も、最初の表示までの時間なのか、使える結果全体が完成するまでなのか、正確に定義する必要があります。

A fast first token can improve perceived responsiveness while leaving the total task slow because of retrieval, tools, or repeated verification.

首个 token 返回得快,可以让人感觉响应更及时,但检索、工具调用或反复验证仍可能使整个任务耗时较长。

最初のトークンが速いと反応は良く感じられますが、検索、ツール、繰り返しの検証で、作業全体は遅いままかもしれません。

Measure ordinary and slow-tail behavior, and include timeout and cancellation paths in the user experience.

应测量常规表现和最慢的一部分请求的表现,并在用户体验中考虑超时与取消流程。

通常時と特に遅い場合を測り、タイムアウトやキャンセルの流れも利用体験に含めます。

Operational controls should restrict data access, protect secrets, and prevent untrusted retrieved text from becoming authority to perform an action.

运行控制措施应限制数据访问、保护密钥等敏感信息,并防止不可信的检索文本被当成执行操作的授权依据。

運用上の制御では、データアクセスの制限、秘密情報の保護、信頼できない取得文書が操作権限として扱われることの防止が必要です。

These controls belong in the system design and cannot be replaced by a polite instruction asking the model to be careful.

这些控制应落实在系统设计中,不能用一句礼貌地要求模型“小心一点”的指令代替。

これらはシステム設計に組み込むもので、モデルに注意するよう丁寧に頼む指示では置き換えられません。

A limited rollout can expose real workflow problems before the system is used for every case.

在系统用于所有案例之前,小范围上线可以暴露真实工作流程中的问题。

限定した導入なら、すべての事例で使用する前に実際の業務上の問題を発見できます。

Define what will be monitored, who can stop the system, and which previous process or version will be restored if performance deteriorates.

应明确监测什么、谁有权停止系统,以及表现恶化时恢复到哪个先前流程或版本。

何を監視し、誰が停止でき、性能が悪化したらどの以前の手順や版へ戻すかを定義します。

A model update, source change, or prompt revision can alter behavior, so version the components that produced each important evaluation result.

模型更新、来源变化或提示词修订都可能改变系统行为,因此应记录产生每项重要评估结果的各个组件版本。

モデルの更新、出典の変更、プロンプト修正で動作が変わるため、重要な評価結果を生んだ構成要素の版を記録します。

A defensible recommendation connects user value, measured quality, operating constraints, and recovery plans, subject to the evidence actually available.

一项有理有据的建议,应把用户价值、测得的质量、运行约束和恢复计划联系起来,并以实际可获得的证据为限。

妥当な推奨は、実際に得られた根拠の範囲内で、利用者の価値、測定した品質、運用制約、復旧計画を結び付けます。

Key terms