16 karma · joined May 11, 2017
Three things bit us: finish_reason semantics differ between "compatible" providers, the model retried identical failing tool calls instead of adapting, and provider failover invalidated prompt caches on both sides.
Curious if others have hit similar issues.
Secrets come from aws secret manager and never injected into env directly.
Each part of the agentic workflow only gets the secrets it needs injected. Agent can see env var names but not the values (our harness masks them) . We also mask any attempts to output to stdout/files.
This keeps the agent architecture simple with env vars that all agents can operate on as it locally. Prompt injection attempts will only yield masked values
Has been working well for us so far
As you add more of them though, they add tech debt and make the code harder to reason about. Developers are rarely motivated to clean them up after rollout.
The main challenge is the increasing complexity of software and processes that necessitate such tools as the engineering team grows. Usually teams building internal tools are seen as cost centers and don't get the same funding as product teams which eventually leads to teams buying more than build to save on future maintenance.
It tends to break apart when working with large teams or when working on mobile apps where one can't just rollback a change easily after it's shipped. The descriptions need to be more detailed and capture lot more information.
For example, our team would require all mobile devs to add information about feature flags for each change to turn the feature off if things go wrong. This also places additional review burden to make sure the flag covers all the new code introduced and does not interact in a bad way with other existing feature flags
Most code review tooling is too coarse in the sense that it expects a small number of reviewers to know the full context of the change or has too many reviewers that slow down things.
Tools to show relevant parts of the change to specific reviewers and potentially break apart a large change into smaller ones automatically will go a long way especially in large scale codebases or when team sizes are large enough that the full context is not understood by everyone on the team
I usually go to the url box to copy the url. one day, it suggested me to use a keyboard shortcut instead based on how often I did it. Little things like this that are built-in to the experience without having to install plugins and make sure they work together is a breath of fresh air.
I have also seen models where the source is made available just before that happens. For eg: https://www.placemark.io/post/placemark-is-winding-down