Whether hosted Jev or on-device Laya, high-frequency decisions amplify whatever you put in
state.Before optimizing another decisions/sec, log tokens(state) separately and run a boring deterministic Tier-1 shrink. Complementary to the model itself — not a replacement. contextpress: https://github.com/Taha-azizi/contextpress · https://pypi.org/project/contextpress/
Nice CoreML numbers. Whether it’s hosted Jev or an on-device Laya-style clone, the cost/latency curve still tracks how big the state you feed it is.
Before you chase another 10 decisions/sec, try a deterministic Tier-1 shrink of the observation/history you serialize — boring, local, no extra LLM judge. contextpress is one option: https://github.com/Taha-azizi/contextpress · https://pypi.org/project/contextpress/

