服务端 agent
Python
唯一能把计算和等待分开的 agent。别的都只会告诉你代码变慢了,然后就没有下文 —— 这一个会说是两者中的哪一个,而它们需要完全相反的修法。
- 包
sixty-sh发布在 PyPI- 运行环境
- Python 3.8 及以上。没有依赖。
- 源码
- sixty-sh/sixty-python
安装
Django, Flask, FastAPI or any WSGI/ASGI app: functions, HTTP routes and SQL, plus the CPU and waiting split only Python can measure.
安装说明是写给你已经开着的那个编码助手的提示词,而不是给你的一份清单。这是有意的:它说的是安装完成之后什么必须成立,而不是要改哪些文件 —— 因为代码该放在哪里取决于框架,而放错地方是会静默失败的。助手可以读你的仓库并把这件事推断出来;文档页上的一个段落做不到。
同样这段文字,就是 install_sixty 通过 MCP 服务器 返回的内容,也是收集端在以下位置提供的内容: /v1/setup?kind=python. 它只有一份。
Python 的完整安装说明
Install the sixty agent in this Python service so its functions, HTTP routes
and database queries report to sixty.
Before editing, inspect whether this service makes LLM or agent calls. If it
does, ask the user: "Do you want LLM monitoring?" If yes, use sixty.agent,
sixty.generation and sixty.tool around custom orchestration, and call
sixty.instrument_openai(client) or sixty.instrument_anthropic(client) for
those SDKs. Never record prompts, outputs, tool arguments or tool results.
1. Add sixty-sh to this project's runtime dependencies — the same place the
web framework is declared (pyproject.toml, requirements.txt, Pipfile), not
a dev or test group. It runs in production; that is the entire point. It
has no dependencies of its own.
2. Call sixty.init() once, as early in process startup as you can get it, and
before any database connection is opened. Put it at the top of the module
the deployed process actually starts — wsgi.py, asgi.py, main.py, manage.py
for a management command — not inside a function that runs per request.
3. Wrap the application so requests become operations. Work out which of these
this project is; do exactly one:
a. Flask — call instrument_flask(app) from "sixty.instrument.flask" after
the app and its routes exist. It wraps app.wsgi_app and names operations
by the matched url_rule.
b. Django — put "sixty.instrument.django.SixtyMiddleware" FIRST in the
MIDDLEWARE list, so the span covers the rest of the middleware rather
than sitting inside it. Operations are named by the route as written in
urls.py.
c. FastAPI, Starlette, Litestar, Quart, or anything else ASGI — wrap with
SixtyASGIMiddleware from "sixty.instrument.asgi".
d. Any other WSGI application — wrap the WSGI callable with SixtyMiddleware
from "sixty.instrument.wsgi".
4. Mark the functions worth measuring. This step is what turns "this endpoint
got slow" into "this function started issuing 14 queries", and skipping it
leaves the feed with routes and queries and nothing in between:
- Put @sixty.trace on the functions that do the work — the service layer,
the repository, whatever this project calls the code between the view and
the database. Not on view functions the middleware already covers.
- Or, for a module of them, call sixty.instrument_module(sys.modules[__name__])
at the bottom of the file; it wraps every public function that module
defines and leaves imported ones alone.
- Leave anything called hundreds of thousands of times a second alone. A
span costs a couple of microseconds, which is nothing next to a request
and everything next to a tight inner loop.
5. Set these environment variables wherever the service is deployed:
SIXTY_API_KEY = a secret key starting sixty_sk_ — ask me for it. Do not
invent one, and do not commit it.
SIXTY_SERVICE = my-app
SIXTY_ENDPOINT = https://ingest.sixty.sh
The release identifier is picked up automatically on Vercel, Render,
Railway, Fly, Heroku and GitHub Actions. If this deploys some other way,
set SIXTY_RELEASE to the commit SHA — without one, every measurement lands
in a single nameless bucket and no comparison can ever be made.
Constraints — correctness requirements, not style preferences:
- Do NOT change any application behaviour. This is instrumentation only: no
refactors, no reordering of business logic, no "while I was in here" fixes.
- Do NOT call init() at import time in a module that is also imported by test
collection or by a build step. Without a key it is inert, but a flush thread
started in a test runner is a surprise nobody asked for.
- Do NOT wrap generators or async generators with @sixty.trace. Their work
happens between next() calls, so the measurement would be of constructing an
object. Wrap whatever drains them.
- Do NOT add any analytics, user id, session id, or cookie to what is
reported. The agent is deliberately anonymous and must stay that way.
- Queries are instrumented through psycopg (2 and 3) automatically. If this
project talks to its database some other way, tell me rather than wiring
something up — measuring it may need work in the agent.
If this service runs under gunicorn, uwsgi or any pre-fork server, note how
many workers it runs: each one reports independently and the collector merges
them, which is correct, but it is worth knowing when you read the numbers.
When you are done, tell me which files you changed and what the deployed start
command now is, so I can confirm data is arriving.它需要一个私密密钥 —— 以 sixty_sk_ 开头,并且只留在服务端。登录之后可以在设置页生成一个。
它测量什么
| 信号 | 单位 | 含义 |
|---|---|---|
rows | rows per call | this query returns more rows than it used to |
fanout | queries per call | this operation now issues more database calls per invocation — an N+1 |
latency | ms per call | this operation takes longer end to end than it used to |
self_latency | ms per call | the time spent in this function itself got longer — its children did not |
payload | bytes per call | the serialized result of this operation got bigger |
errors | error rate | a larger fraction of calls are throwing |
runaway | calls per minute | this operation is being called far more often than anything triggers it |
repeated_query | times per request | the identical query runs several times within one request |
overfetch | rows per call | far more rows are fetched than the code appears to use |
unbounded | rows per call | this query has no upper bound on what it can return |
recursion | levels deep | this operation calls itself, deeper than it should |
new_error | occurrences | an error that did not occur in the previous release |
missing_tenancy | — | This reads a table of per-person data without saying whose rows it wants. Unless your database is filtering it for you, everyone gets everyone else's. |
collapse | — | This is handing back roughly half the data it used to, or less. If that was not deliberate, something is filtering out rows that somebody expects to see. |
vanished | — | It was being used steadily until this release and has not been used once since. Usually the link, button, or redirect that led here stopped working. |
traffic_drop | — | This is still being used, but a fraction as often, and its share of your traffic fell too — so it is not just a quiet period. |
cpu | ms of CPU per call | this function burns more processor time per call than it used to — it is doing more work, not waiting longer |
blocked | ms of waiting per call | this operation spends longer waiting for its turn while doing exactly the same amount of work |
它接在哪里
- Flask — instrument_flask(app) —— 操作按匹配到的 url_rule 命名。
- Django — SixtyMiddleware 放在 MIDDLEWARE 的第一位 —— 按 urls.py 里写的那个路由命名。
- FastAPI、Starlette、Litestar、Quart — SixtyASGIMiddleware,或者任何其他 ASGI 应用。
- 任何 WSGI — Pyramid、Bottle、wsgiref,甚至还没有人写出来的框架。
- 你自己的函数 — 在服务层加 @sixty.trace,或者用 instrument_module() 一次性处理整个模块。
数据库
- psycopg — 2 和 3 两个版本,打在游标上。行数、语句形状,以及到调用函数的归属。
只有它才做的事
- 每次调用的 CPU — time.thread_time() 是按线程算的,读一次约 100 纳秒,而一个同步 span 在整个时长内独占自己的线程 —— 所以这个数字是精确的,而不是摊出来的。
- 每次调用的等待 — 不属于计算的那部分自身时间:一把锁住了更多工作的锁、一个没有空位的连接池、一个握着 GIL 的 C 扩展。这件事发生时其他所有按次统计的数字都还是对的,所以别的东西都抓不到它。
它做不到什么
- 一个跨越 await 的 span 会和事件循环期间跑的其他东西共享线程,所以它上报「没有 CPU」,而不是一个虚高的数字。asyncio 服务能拿到按同步函数和按查询统计的 CPU,但拿不到按请求统计的。
- 不采集查询计划。Postgres 就在那儿,EXPLAIN 也能用;只是 agent 目前还不会发出它。
- 只有 psycopg 被埋了点。架在 psycopg 上的 SQLAlchemy 能测到,因为下面的游标被测到了;asyncpg 和 MySQL 的驱动测不到。
- 在 gunicorn 或 uwsgi 下,每个 worker 各自上报,由收集端合并。这是对的,但当你读一个按进程算的数字时,值得知道这一点。
配置
每个 agent 都读同样四个变量,而且凡是 SIXTY_* 能用的地方 DRIFT_* 依然有效 —— 产品改过名,但那个名字不是我们说撤就能从别人的部署里撤掉的。
SIXTY_API_KEY | 没有它,agent 就保持沉默不动,并且会说出来。它从不猜测,从不对着一个未知端点重试,也从不抛异常。 |
|---|---|
SIXTY_SERVICE | 这个服务叫什么。在能读出项目名的地方,默认用项目名。 |
SIXTY_RELEASE | 最重要的一个。在 Vercel、Render、Railway、Fly、Heroku 和 GitHub Actions 上会自动取到;其他地方请把它设成 commit 的 SHA。没有它,所有测量都会落进同一个没有名字的桶里,任何比较都无从谈起。 |
SIXTY_ENDPOINT | 往哪里上报。默认是 http://localhost:4319,这在笔记本上是对的,而在应用被交付给别人的那一刻就是错的。 |
其余的 —— 发送间隔、采样率、要给什么埋点 —— 都在这个包自己的 README 里,因为那里才是它能随着 agent 变化而保持正确的地方。