Agent framework comparison: what LangChain, CrewAI, AutoGen, and the OpenAI Agents SDK actually lock you into
Every agent framework comparison you’ll find online is a feature table. Does it support tool calling, does it support memory, does it support multi-agent handoff. Yes, they all do, more or less. That’s not the question that matters once you’ve shipped something and it’s handling real traffic with a real API bill attached to it. The question that matters is: what does this framework own, and what happens to your code the day you want to swap it out.
I’ve built on top of a few of these now, for real workloads, not demos. The pattern I keep running into is that lock-in doesn’t show up in the docs. It shows up in where your state lives, whose object model your business logic gets written against, and whether your orchestration logic can survive being pulled out of the framework at all. Here’s how the major ones actually hold onto you.
What “lock-in” means for an agent framework specifically
Three things create exit cost, and they’re not the same thing:
State storage. If your conversation history, tool call results, and intermediate reasoning steps live in a database or object format that only that framework’s runtime knows how to read, you can’t leave without a migration project.
Orchestration model. Some frameworks make you express your agent logic as their graph, their group chat, or their crew of role-playing agents. That logic is now written in their vocabulary, not yours, and porting it means re-deriving the control flow, not just swapping an SDK call.
Tooling gravity. Tracing, evals, and debugging tools that are tied to one framework’s execution model pull you deeper in even after the core logic is portable, because you don’t want to lose your observability when you leave.
Most comparisons only look at the first one. All three matter, and the second one is usually the expensive one to unwind.
LangChain and LangGraph: portable core, sticky graph
LangChain’s chain and LCEL layer is genuinely loosely coupled. A chain is close to a plain function composition, and you can lift most of that logic out without much pain, since it’s built to wrap calls to whatever model provider you point it at.
LangGraph is a different story. It models your agent as a state machine with nodes and edges, and it persists that state through its own checkpointer objects. Once your agent’s control flow (should it loop, should it hand off, should it wait for a human) is expressed as a LangGraph graph, that control flow is written in LangGraph’s terms. You can extract the individual node functions, since those are usually just Python calling an LLM or a tool. What you can’t cleanly extract is the graph topology and the checkpointing behavior, because that’s the part LangGraph actually owns. If you’re also using LangSmith for tracing, your observability is tied to the same ecosystem, so leaving means rebuilding your eval and debug workflow somewhere else at the same time you’re rebuilding your orchestration.
OpenAI’s Assistants API and Agents SDK: state lives on their servers
The old Assistants API was the clearest lock-in case in the whole space, because it wasn’t a client-side library decision at all, it was an infrastructure decision. Threads, runs, and messages were stored server-side on OpenAI’s infrastructure. Your app didn’t hold the conversation state, OpenAI did. That’s convenient right up until you want to switch model providers for cost or latency reasons, at which point you discover your entire conversation history and file search indexes live somewhere you can’t point at a different vendor. OpenAI has been moving this surface toward the Responses API and the open-source Agents SDK, which pushes more of the orchestration logic into code you control rather than a server-side thread object. That’s a meaningfully different lock-in profile: closer to a normal SDK, less of a “your data lives on our servers in our format” problem. But the tool-calling schema and function definitions are still written against OpenAI’s expected shapes, so swapping the underlying model provider still means touching every tool definition, even if your conversation state is now something you own.
CrewAI: the role-playing abstraction becomes your architecture
CrewAI’s whole value proposition is that you describe agents as roles with goals and backstories, group them into a crew, and let a process (sequential or hierarchical) decide how tasks flow between them. That’s a genuinely different way to think about agent design than a raw function-calling loop, and for some multi-step workflows it maps well onto how you’d naturally divide the work between specialists.
The lock-in cost is that your task decomposition ends up expressed as CrewAI’s Task and Crew objects, and its more recent YAML-based configuration format ties the definition of your agents directly to CrewAI’s schema rather than to your own code. If you leave, you’re not just swapping an LLM call, you’re re-deriving why the work was split the way it was, because that reasoning was implicit in how you defined the crew, not written out anywhere else.
AutoGen and AG2: the conversation is the API
AutoGen (and its community fork AG2, after part of the original team split off) models multi-agent systems as agents talking to each other in a shared conversation, coordinated by a group chat manager. It’s a good fit for problems that are naturally dialogic, a critic agent reviewing a writer agent’s output, for example.
The catch is that the message-passing pattern is the architecture. Your business logic is encoded in who’s allowed to speak next and what triggers a handoff, which lives inside AutoGen’s group chat orchestration, not in a control flow you wrote yourself. Extracting that into a different framework means re-implementing the turn-taking logic from scratch, because there’s no clean seam between “the conversation” and “the app.”
Semantic Kernel: plugins and planners, Microsoft’s vocabulary
Semantic Kernel takes the plugin approach, formerly called skills, where your tools and functions are registered with a kernel object and invoked through a planner. It’s built to run across C#, Python, and Java, which is a real advantage if your org already has a polyglot backend and doesn’t want to standardize on Python just to get agent tooling.
The lock-in here is in the planner and the kernel’s own memory and connector abstractions. Once your tool functions are decorated and registered as Semantic Kernel plugins, and your planning logic depends on the kernel’s function-calling and memory retrieval conventions, unwinding that means rewriting the plugin registration layer and, if you used its planners, rebuilding the reasoning strategy that decided which plugin to call when.
The pattern underneath all of them
Every one of these frameworks sells you the same thing: a higher-level vocabulary for describing agent behavior so you write less boilerplate. And every one of them collects the cost of that convenience in the same place, the part of your system where control flow decisions get made. The tool-calling glue is usually the least sticky part, since most frameworks are just wrapping a chat completions style API underneath. The orchestration layer, the part that decides what happens next, is where you actually get locked in, because that’s the part that only makes sense inside that framework’s model of how agents behave.
What actually determines your exit cost
If I’m evaluating one of these for a new project, I ask three questions before I ask about features. Where does session state actually get persisted, and can I read it with a plain database query if I had to. Can I extract my control flow logic as functions I own, or is it expressed entirely as configuration objects belonging to the framework. And does my tracing and debugging setup survive if I strip the framework out, or does leaving mean flying blind on observability while I rebuild.
None of these frameworks are wrong choices. They’re all solving a real problem, and the convenience is genuine, not marketing. But “which one is more powerful” is the wrong first question. The first question is which parts of your system you’re willing to write in someone else’s vocabulary, because that’s the part you’ll be rewriting the day you outgrow them.
If you want more of this kind of breakdown on the tools you’re actually building with, come find us at AI Tool Gazette.