diff --git a/docs/docs/cloud/concepts/api.md b/docs/docs/cloud/concepts/api.md index f8e948285..7b2f56b2d 100644 --- a/docs/docs/cloud/concepts/api.md +++ b/docs/docs/cloud/concepts/api.md @@ -62,122 +62,7 @@ The LangGraph Cloud API offers several features to support complex agent archite ### Streaming -Streaming is critical for making LLM applications feel responsive to end users. When creating a streaming run, the streaming mode determines what data is streamed back to the API client. The LangGraph Cloud API supports five streaming modes. - -- `values`: Stream the full state of the graph after each [super-step](https://langchain-ai.github.io/langgraph/concepts/low_level/#graphs) is executed. See the [how-to guide](../how-tos/stream_values.md) for streaming values. -- `messages`: Stream complete messages (at the end of node execution) as well as tokens for any messages generated inside a node. This mode is primarily meant for powering chat applications. This is only an option if your graph contains a `messages` key. See the [how-to guide](../how-tos/stream_messages.md) for streaming messages. -- `updates`: Streams updates to the state of the graph after each node is executed. See the [how-to guide](../how-tos/stream_updates.md) for streaming updates. -- `events`: Stream all events (including the state of the graph) that occur during graph execution. See the [how-to guide](../how-tos/stream_events.md) for streaming events. This can be used to do token-by-token streaming for LLMs. -- `debug`: Stream debug events throughout graph execution. See the [how-to guide](../how-tos/stream_debug.md) for streaming debug events. - -You can also specify multiple streaming modes at the same time. See the [how-to guide](../how-tos/stream_multiple.md) for configuring multiple streaming modes at the same time. - -See the [API reference](../reference/api/api_ref.html#tag/runscreate/POST/threads/{thread_id}/runs/stream) for how to create streaming runs. - -Streaming modes `values`, `updates`, and `debug` are very similar to modes available in the LangGraph library - for a deeper conceptual explanation of those, you can see the LangGraph library documentation [here](../../concepts/low_level.md#streaming). - -Streaming mode `events` is the same as using `.astream_events` in the LangGraph library - for a deeper conceptual explanation of this, you can see the LangGraph library documentation [here](../../concepts/low_level.md#streaming). - -#### `mode="messages"` -Streaming mode `messages` is a new streaming mode, currently only available in the API. What does this mode enable? - -This mode is focused on streaming back messages. It currently assumes that you have a `messages` key in your graph that is a list of messages. Assuming we have a simple react agent deployed, what does this stream look like? - -All events emitted have two attributes: - -- `event`: This is the name of the event -- `data`: This is data associated with the event - -Let's run it on a question that should trigger a tool call: - -```python -thread = await client.threads.create() -input = {"messages": [{"role": "user", "content": "what's the weather in sf?"}]} - -events = [] -async for event in client.runs.stream( - thread["thread_id"], - assistant_id="agent", # This may need to change depending on the graph you deployed - input=input, - stream_mode="messages", -): - print(event.event) -``` -```shell -metadata -messages/complete -messages/metadata -messages/partial -... -messages/partial -messages/complete -messages/complete -messages/metadata -messages/partial -... -messages/partial -messages/complete -end -``` - -We first get some `metadata` - this is metadata about the run. - -```python -StreamPart(event='metadata', data={'run_id': '1ef657cf-ae55-6f65-97d4-f4ed1dbdabc6'}) -``` - -We then get a `messages/complete` event - this a fully formed message getting emitted. In this case, -this was the just the input message we sent in. - -```python -StreamPart(event='messages/complete', data=[{'content': 'hi!', 'additional_kwargs': {}, 'response_metadata': {}, 'type': 'human', 'name': None, 'id': '833c09a3-bb19-46c9-81d9-1e5954ec5f92', 'example': False}]) -``` - -We then get a `messages/metadata` - this is just letting us know that a new message is starting. - -```python -StreamPart(event='messages/metadata', data={'run-985c0f14-9f43-40d4-a505-4637fc58e333': {'metadata': {'created_by': 'system', 'run_id': '1ef657de-7594-66df-8eb2-31518e4a1ee2', 'graph_id': 'agent', 'thread_id': 'c178eab5-e293-423c-8e7d-1d113ffe7cd9', 'model_name': 'openai', 'assistant_id': 'fe096781-5601-53d2-b2f6-0d3403f7e9ca', 'langgraph_step': 1, 'langgraph_node': 'agent', 'langgraph_triggers': ['start:agent'], 'langgraph_task_idx': 0, 'ls_provider': 'openai', 'ls_model_name': 'gpt-4o', 'ls_model_type': 'chat', 'ls_temperature': 0.0}}}) -``` - -We then get a BUNCH of `messages/partial` events - these are the individual tokens from the LLM! In the case below, we can see the START of a tool call. - -```python -StreamPart(event='messages/partial', data=[{'content': '', 'additional_kwargs': {'tool_calls': [{'index': 0, 'id': 'call_w8Hr8dHGuZCPgRfd5FqRBArs', 'function': {'arguments': '', 'name': 'tavily_search_results_json'}, 'type': 'function'}]}, 'response_metadata': {}, 'type': 'ai', 'name': None, 'id': 'run-985c0f14-9f43-40d4-a505-4637fc58e333', 'example': False, 'tool_calls': [], 'invalid_tool_calls': [{'name': 'tavily_search_results_json', 'args': '', 'id': 'call_w8Hr8dHGuZCPgRfd5FqRBArs', 'error': None}], 'usage_metadata': None}]) -``` - -After that, we get a `messages/complete` event - this is the AIMessage finishing. It's now a complete tool call: - -```python -StreamPart(event='messages/complete', data=[{'content': '', 'additional_kwargs': {'tool_calls': [{'index': 0, 'id': 'call_w8Hr8dHGuZCPgRfd5FqRBArs', 'function': {'arguments': '{"query":"current weather in San Francisco"}', 'name': 'tavily_search_results_json'}, 'type': 'function'}]}, 'response_metadata': {'finish_reason': 'tool_calls', 'model_name': 'gpt-4o-2024-05-13', 'system_fingerprint': 'fp_157b3831f5'}, 'type': 'ai', 'name': None, 'id': 'run-985c0f14-9f43-40d4-a505-4637fc58e333', 'example': False, 'tool_calls': [{'name': 'tavily_search_results_json', 'args': {'query': 'current weather in San Francisco'}, 'id': 'call_w8Hr8dHGuZCPgRfd5FqRBArs'}], 'invalid_tool_calls': [], 'usage_metadata': None}]) -``` - -After that, we get ANOTHER `messages/complete` event. This is a tool message - our agent has called a tool, gotten a response, and now inserting it into the state in the form of a tool message. - -```python -StreamPart(event='messages/complete', data=[{'content': '[{"url": "https://www.weatherapi.com/", "content": "{\'location\': {\'name\': \'San Francisco\', \'region\': \'California\', \'country\': \'United States of America\', \'lat\': 37.78, \'lon\': -122.42, \'tz_id\': \'America/Los_Angeles\', \'localtime_epoch\': 1724877689, \'localtime\': \'2024-08-28 13:41\'}, \'current\': {\'last_updated_epoch\': 1724877000, \'last_updated\': \'2024-08-28 13:30\', \'temp_c\': 23.3, \'temp_f\': 73.9, \'is_day\': 1, \'condition\': {\'text\': \'Partly cloudy\', \'icon\': \'//cdn.weatherapi.com/weather/64x64/day/116.png\', \'code\': 1003}, \'wind_mph\': 15.0, \'wind_kph\': 24.1, \'wind_degree\': 310, \'wind_dir\': \'NW\', \'pressure_mb\': 1014.0, \'pressure_in\': 29.93, \'precip_mm\': 0.0, \'precip_in\': 0.0, \'humidity\': 57, \'cloud\': 25, \'feelslike_c\': 25.0, \'feelslike_f\': 77.1, \'windchill_c\': 20.9, \'windchill_f\': 69.6, \'heatindex_c\': 23.3, \'heatindex_f\': 74.0, \'dewpoint_c\': 12.9, \'dewpoint_f\': 55.2, \'vis_km\': 16.0, \'vis_miles\': 9.0, \'uv\': 6.0, \'gust_mph\': 19.5, \'gust_kph\': 31.3}}"}]', 'additional_kwargs': {}, 'response_metadata': {}, 'type': 'tool', 'name': 'tavily_search_results_json', 'id': '0112eba5-7660-4375-9f24-c7a1d6777b97', 'tool_call_id': 'call_w8Hr8dHGuZCPgRfd5FqRBArs'}]) -``` - -After that, we see the agent doing another LLM call and streaming back a response. We then get an `end` event: - -```python -StreamPart(event='end', data=None) -``` - -And that's it! This is more focused streaming mode specifically focused on streaming back messages. See this [how-to guide](../how-tos/stream_messages.md) for more information. - - -### Human-in-the-Loop - -There are many occasions where the graph cannot run completely autonomously. For instance, the user might need to input some additional arguments to a function call, or select the next edge for the graph to continue on. In these instances, we need to insert some human in the loop interaction, which you can learn about in the [human in the loop how-tos](../how-tos/index.md#human-in-the-loop). - -### Double Texting - -Many times users might interact with your graph in unintended ways. For instance, a user may send one message and before the graph has finished running send a second message. To solve this issue of "double-texting" (i.e. prompting the graph a second time before the first run has finished), LangGraph has provided four different solutions, all of which are covered in the [Double Texting how-tos](../how-tos/index.md#double-texting). These options are: - -- `reject`: This is the simplest option, this just rejects any follow up runs and does not allow double texting. See the [how-to guide](../how-tos/reject_concurrent.md) for configuring the reject double text option. -- `enqueue`: This is a relatively simple option which continues the first run until it completes the whole run, then sends the new input as a separate run. See the [how-to guide](../how-tos/enqueue_concurrent.md) for configuring the enqueue double text option. -- `interrupt`: This option interrupts the current execution but saves all the work done up until that point. It then inserts the user input and continues from there. If you enable this option, your graph should be able to handle weird edge cases that may arise. See the [how-to guide](../how-tos/interrupt_concurrent.md) for configuring the interrupt double text option. -- `rollback`: This option rolls back all work done up until that point. It then sends the user input in, basically as if it just followed the original run input. See the [how-to guide](../how-tos/rollback_concurrent.md) for configuring the rollback double text option. +[Streaming](../../concepts/streaming.md) is critical for making LLM applications feel responsive to end users. When creating a streaming run, the streaming mode determines what data is streamed back to the API client. LangGraph Platform supports five streaming modes: `values`, `updates`, `messages-tuple`, `events`, and `debug`. See these [how-to guides](../../how-tos/index.md#streaming_1) and the [API reference](../reference/api/api_ref.html#tag/thread-runs/POST/threads/%7Bthread_id%7D/runs/stream) for more details. ### Stateless Runs diff --git a/docs/docs/cloud/how-tos/stream_messages.md b/docs/docs/cloud/how-tos/stream_messages.md index d81d40a57..a359956a3 100644 --- a/docs/docs/cloud/how-tos/stream_messages.md +++ b/docs/docs/cloud/how-tos/stream_messages.md @@ -3,9 +3,7 @@ !!! info "Prerequisites" * [Streaming](../../concepts/streaming.md) -This guide covers how to stream messages from your graph. With `stream_mode="messages"`, messages from any chat model invocations inside your graph nodes will be streamed back. - -Read more about how the `messages` streaming mode works [here](https://langchain-ai.github.io/langgraph/cloud/concepts/api/#modemessages) +This guide covers how to stream messages from your graph. With `stream_mode="messages-tuple"`, messages (i.e. individual LLM tokens) from any chat model invocations inside your graph nodes will be streamed back. ## Setup @@ -60,7 +58,7 @@ Output: ## Stream graph in messages mode -Now we can stream by messages, which will return complete messages (at the end of node execution) as well as tokens for any messages generated inside a node: +Now we can stream LLM tokens for any messages generated inside a node in the form of tuples `(message, metadata)`. Metadata contains additional information that can be useful for filtering the streamed outputs to a specific node or LLM. === "Python" @@ -73,7 +71,7 @@ Now we can stream by messages, which will return complete messages (at the end o assistant_id=assistant_id, input=input, config=config, - stream_mode="messages", + stream_mode="messages-tuple", ): print(f"Receiving new event of type: {chunk.event}...") print(chunk.data) @@ -99,7 +97,7 @@ Now we can stream by messages, which will return complete messages (at the end o { input, config, - streamMode: "messages" + streamMode: "messages-tuple" } ); for await (const chunk of streamResponse) { @@ -119,7 +117,7 @@ Now we can stream by messages, which will return complete messages (at the end o \"assistant_id\": \"agent\", \"input\": {\"messages\": [{\"role\": \"human\", \"content\": \"what's the weather in la\"}]}, \"stream_mode\": [ - \"messages\" + \"messages-tuple\" ] }" | \ sed 's/\r$//' | \ @@ -150,117 +148,101 @@ Output: Receiving new event of type: metadata... {"run_id": "1ef971e0-9a84-6154-9047-247b4ce89c4d", "attempt": 1} - - - Receiving new event of type: messages/metadata... - { - "run-700157a5-df1a-4829-9e7c-1e07a1d934f7": { - "metadata": { - "graph_id": "agent", - "langgraph_node": "agent", - ... - } - } - } - ... - Receiving new event of type: messages/partial... + Receiving new event of type: messages... [ { + "type": "AIMessageChunk", "tool_calls": [ { "name": "tavily_search_results_json", "args": { - "query": "weather" + "query": "weat" }, - "id": "toolu_01RJGmVJtTxccoHHixGkGqaC", - "type": "tool_call" - } - ], - } - ] - - - - Receiving new event of type: messages/partial... - [ - { - "type": "ai", - "tool_calls": [ - { - "name": "tavily_search_results_json", - "args": { - "query": "weather in " - }, - "id": "toolu_01RJGmVJtTxccoHHixGkGqaC", + "id": "toolu_0114XKXdNtHQEa3ozmY1uDdM", "type": "tool_call" } ], ... - } - ] - - ... - - Receiving new event of type: messages/partial... - [ + }, { - "type": "ai", - "tool_calls": [ - { - "name": "tavily_search_results_json", - "args": { - "query": "weather in san francisco" - }, - "id": "toolu_01RJGmVJtTxccoHHixGkGqaC", - "type": "tool_call" - } - ], + "graph_id": "agent", + "langgraph_node": "agent", ... } ] - Receiving new event of type: messages/metadata... - { - "aa162b98-433d-4e3c-b204-0d41a6694156": { - "metadata": { - "graph_id": "agent", - "langgraph_node": "action", - ... - } - } - } - - - - Receiving new event of type: messages/complete... + Receiving new event of type: messages... [ { - "content": "[{\"url\": \"https://www.weatherapi.com/\", \"content\": \"{'location': {'name': 'San Francisco', 'region': 'California', 'country': 'United States of America', 'lat': 37.775, 'lon': -122.4183, 'tz_id': 'America/Los_Angeles', 'localtime_epoch': 1730334046, 'localtime': '2024-10-30 17:20'}, 'current': {'last_updated_epoch': 1730333700, 'last_updated': '2024-10-30 17:15', 'temp_c': 12.3, 'temp_f': 54.2, 'is_day': 1, 'condition': {'text': 'Partly Cloudy', 'icon': '//cdn.weatherapi.com/weather/64x64/day/116.png', 'code': 1003}, 'wind_mph': 9.6, 'wind_kph': 15.5, 'wind_degree': 238, 'wind_dir': 'WSW', 'pressure_mb': 1021.0, 'pressure_in': 30.15, 'precip_mm': 0.0, 'precip_in': 0.0, 'humidity': 93, 'cloud': 57, 'feelslike_c': 11.2, 'feelslike_f': 52.2, 'windchill_c': 11.2, 'windchill_f': 52.2, 'heatindex_c': 12.3, 'heatindex_f': 54.2, 'dewpoint_c': 11.2, 'dewpoint_f': 52.1, 'vis_km': 10.0, 'vis_miles': 6.0, 'uv': 0.5, 'gust_mph': 12.9, 'gust_kph': 20.8}}\"}]", + "type": "AIMessageChunk", + "tool_calls": [ + { + "name": "tavily_search_results_json", + "args": { + "query": "her in san " + }, + "id": "toolu_0114XKXdNtHQEa3ozmY1uDdM", + "type": "tool_call" + } + ], + ... + }, + { + "graph_id": "agent", + "langgraph_node": "agent", + ... + } + ] + + ... + + Receiving new event of type: messages... + [ + { + "type": "AIMessageChunk", + "tool_calls": [ + { + "name": "tavily_search_results_json", + "args": { + "query": "francisco" + }, + "id": "toolu_0114XKXdNtHQEa3ozmY1uDdM", + "type": "tool_call" + } + ], + ... + }, + { + "graph_id": "agent", + "langgraph_node": "agent", + ... + } + ] + + ... + + Receiving new event of type: messages... + [ + { + "content": "[{\"url\": \"https://www.weatherapi.com/\", \"content\": \"{'location': {'name': 'San Francisco', 'region': 'California', 'country': 'United States of America', 'lat': 37.775, 'lon': -122.4183, 'tz_id': 'America/Los_Angeles', 'localtime_epoch': 1730475777, 'localtime': '2024-11-01 08:42'}, 'current': {'last_updated_epoch': 1730475000, 'last_updated': '2024-11-01 08:30', 'temp_c': 11.1, 'temp_f': 52.0, 'is_day': 1, 'condition': {'text': 'Partly cloudy', 'icon': '//cdn.weatherapi.com/weather/64x64/day/116.png', 'code': 1003}, 'wind_mph': 2.2, 'wind_kph': 3.6, 'wind_degree': 192, 'wind_dir': 'SSW', 'pressure_mb': 1018.0, 'pressure_in': 30.07, 'precip_mm': 0.0, 'precip_in': 0.0, 'humidity': 89, 'cloud': 75, 'feelslike_c': 11.5, 'feelslike_f': 52.6, 'windchill_c': 10.0, 'windchill_f': 50.1, 'heatindex_c': 10.4, 'heatindex_f': 50.7, 'dewpoint_c': 9.1, 'dewpoint_f': 48.5, 'vis_km': 16.0, 'vis_miles': 9.0, 'uv': 3.0, 'gust_mph': 6.7, 'gust_kph': 10.8}}\"}]", "type": "tool", - "name": "tavily_search_results_json", - "tool_call_id": "toolu_01RJGmVJtTxccoHHixGkGqaC", + "tool_call_id": "toolu_0114XKXdNtHQEa3ozmY1uDdM", + ... + }, + { + "graph_id": "agent", + "langgraph_node": "action", + ... } ] + ... - Receiving new event of type: messages/metadata... - { - "run-f92646d2-6b13-4648-90c7-0280766bfaf2": { - "metadata": { - "graph_id": "agent", - "langgraph_node": "agent", - ... - } - } - } - - - - Receiving new event of type: messages/partial... + Receiving new event of type: messages... [ { "content": [ @@ -270,41 +252,80 @@ Output: "index": 0 } ], - "type": "ai", + "type": "AIMessageChunk", + ... + }, + { + "graph_id": "agent", + "langgraph_node": "agent", ... } ] - Receiving new event of type: messages/partial... + Receiving new event of type: messages... [ { "content": [ { - "text": "\n\nThe search results provide", + "text": " results provide", "type": "text", "index": 0 } ], - "type": "ai", + "type": "AIMessageChunk", + ... + }, + { + "graph_id": "agent", + "langgraph_node": "agent", ... } ] - ... - Receiving new event of type: messages/partial... + + Receiving new event of type: messages... [ { "content": [ { - "text": "\n\nThe search results provide the current weather conditions in San Francisco. According to the data, as of 5:20pm on October 30, 2024, the weather in San Francisco is partly cloudy with a temperature of 54\\u00b0F (12\\u00b0C). The wind is blowing from the west-southwest at around 10 mph (15 km/h). The humidity is high at 93% and visibility is 6 miles (10 km). Overall, it seems to be a cool, partly cloudy day with moderate winds in San Francisco.", + "text": " the current weather conditions", "type": "text", "index": 0 } ], - "type": "ai", + "type": "AIMessageChunk", + ... + }, + { + "graph_id": "agent", + "langgraph_node": "agent", ... } - ] \ No newline at end of file + ] + + + + Receiving new event of type: messages... + [ + { + "content": [ + { + "text": " in San Francisco.", + "type": "text", + "index": 0 + } + ], + "type": "AIMessageChunk", + ... + }, + { + "graph_id": "agent", + "langgraph_node": "agent", + ... + } + ] + + ... \ No newline at end of file diff --git a/docs/docs/concepts/streaming.md b/docs/docs/concepts/streaming.md index 052715a01..4cff01497 100644 --- a/docs/docs/concepts/streaming.md +++ b/docs/docs/concepts/streaming.md @@ -153,7 +153,7 @@ guide for that [here](../how-tos/streaming-tokens.ipynb). Streaming is critical for making LLM applications feel responsive to end users. When creating a streaming run, the streaming mode determines what data is streamed back to the API client. LangGraph Platform supports five streaming modes: - `values`: Stream the full state of the graph after each [super-step](https://langchain-ai.github.io/langgraph/concepts/low_level/#graphs) is executed. See the [how-to guide](../cloud/how-tos/stream_values.md) for streaming values. -- `messages`: Stream complete messages (at the end of node execution) as well as tokens for any messages generated inside a node. This mode is primarily meant for powering chat applications. This is only an option if your graph contains a `messages` key. See the [how-to guide](../cloud/how-tos/stream_messages.md) for streaming messages. +- `messages-tuple`: Stream LLM tokens for any messages generated inside a node. This mode is primarily meant for powering chat applications. See the [how-to guide](../cloud/how-tos/stream_messages.md) for streaming messages. - `updates`: Streams updates to the state of the graph after each node is executed. See the [how-to guide](../cloud/how-tos/stream_updates.md) for streaming updates. - `events`: Stream all events (including the state of the graph) that occur during graph execution. See the [how-to guide](../cloud/how-tos/stream_events.md) for streaming events. This can be used to do token-by-token streaming for LLMs. - `debug`: Stream debug events throughout graph execution. See the [how-to guide](../cloud/how-tos/stream_debug.md) for streaming debug events. @@ -162,90 +162,11 @@ You can also specify multiple streaming modes at the same time. See the [how-to See the [API reference](../cloud/reference/api/api_ref.html#tag/threads-runs/POST/threads/{thread_id}/runs/stream) for how to create streaming runs. -Streaming modes `values`, `updates`, and `debug` are very similar to modes available in the LangGraph library - for a deeper conceptual explanation of those, you can see the [previous section](#streaming-graph-outputs-stream-and-astream). +Streaming modes `values`, `updates`, `messages-tuple` and `debug` are very similar to modes available in the LangGraph library - for a deeper conceptual explanation of those, you can see the [previous section](#streaming-graph-outputs-stream-and-astream). Streaming mode `events` is the same as using `.astream_events` in the LangGraph library - for a deeper conceptual explanation of this, you can see the [previous section](#streaming-graph-outputs-stream-and-astream). -### `stream_mode="messages"` - -Streaming mode `messages` is for streaming back messages from the LLM. Assuming we have a simple [ReAct](./agentic_concepts.md#react-implementation)-style agent deployed, what does this stream look like? - All events emitted have two attributes: - `event`: This is the name of the event -- `data`: This is data associated with the event - -!!! note - Streaming mode `messages` is different from the one in the LangGraph library: - - - LangGraph Server streams event objects with messages in the `data` field, while LangGraph library streams tuples (`AIMessageChunk`, metadata). - - In LangGraph Server, metadata is streamed only once per message (`messages/metadata`), before the individual tokens are streamed (`messages/partial`), while in LangGraph library it's streamed with every `AIMessageChunk` (for each LLM token). - - LangGraph Server also streams additional events (`metadata`, `messages/complete`, see below for more details). - -Let's run it on a question that should trigger a tool call: - -```python -thread = await client.threads.create() -input = {"messages": [{"role": "user", "content": "what's the weather in sf?"}]} - -events = [] -async for event in client.runs.stream( - thread["thread_id"], - assistant_id="agent", # This may need to change depending on the graph you deployed - input=input, - stream_mode="messages", -): - print(event.event) -``` -```shell -metadata -messages/metadata -messages/partial -... -messages/partial -messages/metadata -messages/complete -messages/metadata -messages/partial -... -messages/partial -end -``` - -We first get some `metadata` - this is metadata about the run. - -```python -StreamPart(event='metadata', data={'run_id': '1ef657cf-ae55-6f65-97d4-f4ed1dbdabc6'}) -``` - -We then get a `messages/metadata` - this is letting us know that a new message is starting and provides additional information about the LLM as well as the node where the LLM is invoked. - -```python -StreamPart(event='messages/metadata', data={'run-985c0f14-9f43-40d4-a505-4637fc58e333': {'metadata': {'created_by': 'system', 'run_id': '1ef657de-7594-66df-8eb2-31518e4a1ee2', 'graph_id': 'agent', 'thread_id': 'c178eab5-e293-423c-8e7d-1d113ffe7cd9', 'model_name': 'openai', 'assistant_id': 'fe096781-5601-53d2-b2f6-0d3403f7e9ca', 'langgraph_step': 1, 'langgraph_node': 'agent', 'langgraph_triggers': ['start:agent'], 'langgraph_task_idx': 0, 'ls_provider': 'openai', 'ls_model_name': 'gpt-4o', 'ls_model_type': 'chat', 'ls_temperature': 0.0}}}) -``` - -We then get a BUNCH of `messages/partial` events - these are the individual tokens from the LLM! In the case below, we can see the START of a tool call. - -```python -StreamPart(event='messages/partial', data=[{'content': '', 'additional_kwargs': {'tool_calls': [{'index': 0, 'id': 'call_w8Hr8dHGuZCPgRfd5FqRBArs', 'function': {'arguments': '', 'name': 'tavily_search_results_json'}, 'type': 'function'}]}, 'response_metadata': {}, 'type': 'ai', 'name': None, 'id': 'run-985c0f14-9f43-40d4-a505-4637fc58e333', 'example': False, 'tool_calls': [], 'invalid_tool_calls': [{'name': 'tavily_search_results_json', 'args': '', 'id': 'call_w8Hr8dHGuZCPgRfd5FqRBArs', 'error': None}], 'usage_metadata': None}]) -``` - -The last `messages/partial` event for a given message will contain all of the tokens streamed for that message. In our case, it is now a complete tool call: - -```python -StreamPart(event='messages/partial', data=[{'content': '', 'additional_kwargs': {'tool_calls': [{'index': 0, 'id': 'call_w8Hr8dHGuZCPgRfd5FqRBArs', 'function': {'arguments': '{"query":"current weather in San Francisco"}', 'name': 'tavily_search_results_json'}, 'type': 'function'}]}, 'response_metadata': {'finish_reason': 'tool_calls', 'model_name': 'gpt-4o-2024-05-13', 'system_fingerprint': 'fp_157b3831f5'}, 'type': 'ai', 'name': None, 'id': 'run-985c0f14-9f43-40d4-a505-4637fc58e333', 'example': False, 'tool_calls': [{'name': 'tavily_search_results_json', 'args': {'query': 'current weather in San Francisco'}, 'id': 'call_w8Hr8dHGuZCPgRfd5FqRBArs'}], 'invalid_tool_calls': [], 'usage_metadata': None}]) -``` - -After that, we get another `messages/metadata`, now followed by a `messages/complete` event. This event is emitted for a tool message - our agent has called a tool, gotten a response, and now inserting it into the state in the form of a tool message. - -```python -StreamPart(event='messages/complete', data=[{'content': '[{"url": "https://www.weatherapi.com/", "content": "{\'location\': {\'name\': \'San Francisco\', \'region\': \'California\', \'country\': \'United States of America\', \'lat\': 37.78, \'lon\': -122.42, \'tz_id\': \'America/Los_Angeles\', \'localtime_epoch\': 1724877689, \'localtime\': \'2024-08-28 13:41\'}, \'current\': {\'last_updated_epoch\': 1724877000, \'last_updated\': \'2024-08-28 13:30\', \'temp_c\': 23.3, \'temp_f\': 73.9, \'is_day\': 1, \'condition\': {\'text\': \'Partly cloudy\', \'icon\': \'//cdn.weatherapi.com/weather/64x64/day/116.png\', \'code\': 1003}, \'wind_mph\': 15.0, \'wind_kph\': 24.1, \'wind_degree\': 310, \'wind_dir\': \'NW\', \'pressure_mb\': 1014.0, \'pressure_in\': 29.93, \'precip_mm\': 0.0, \'precip_in\': 0.0, \'humidity\': 57, \'cloud\': 25, \'feelslike_c\': 25.0, \'feelslike_f\': 77.1, \'windchill_c\': 20.9, \'windchill_f\': 69.6, \'heatindex_c\': 23.3, \'heatindex_f\': 74.0, \'dewpoint_c\': 12.9, \'dewpoint_f\': 55.2, \'vis_km\': 16.0, \'vis_miles\': 9.0, \'uv\': 6.0, \'gust_mph\': 19.5, \'gust_kph\': 31.3}}"}]', 'additional_kwargs': {}, 'response_metadata': {}, 'type': 'tool', 'name': 'tavily_search_results_json', 'id': '0112eba5-7660-4375-9f24-c7a1d6777b97', 'tool_call_id': 'call_w8Hr8dHGuZCPgRfd5FqRBArs'}]) -``` - -After that, we see the agent doing another LLM call and streaming back a response. We then get an `end` event: - -```python -StreamPart(event='end', data=None) -``` - -And that's it! This is more focused streaming mode specifically focused on streaming back messages. See this [how-to guide](../cloud/how-tos/stream_messages.md) for more information. \ No newline at end of file +- `data`: This is data associated with the event \ No newline at end of file diff --git a/docs/mkdocs.yml b/docs/mkdocs.yml index e2813b141..08f15d8bd 100644 --- a/docs/mkdocs.yml +++ b/docs/mkdocs.yml @@ -54,6 +54,9 @@ plugins: - search: separator: '[\s\u200b\-_,:!=\[\]()"`/]+|\.(?!\d)|&[lg]t;|(?!\b)(?=[A-Z][a-z])' - autorefs + - redirects: + redirect_maps: + 'cloud/concepts/api.md': 'concepts/index.md#langgraph-platform' - mkdocstrings: handlers: python: