Lists That Keep Growing
We embedded a cart’s items and an order’s items partly because we expect those lists to stay small. An order records one purchase. Its list of items does not keep growing as the shopper makes more purchases; each purchase creates another order. To see why growth matters, let’s look at a different application: a chat app.
Messages inside a conversation
A conversation has participants and messages. Each message belongs to one conversation, and the app displays messages together. Those are reasons to consider embedding the messages in the conversation document:
{
"_id": "conversation-12",
"participantIds": ["user-7", "user-18"],
"messages": [
{
"_id": "message-101",
"senderId": "user-7",
"sentAt": "2026-09-28T14:00:00Z",
"text": "Are we meeting at three?"
},
{
"_id": "message-102",
"senderId": "user-18",
"sentAt": "2026-09-28T14:01:00Z",
"text": "Yes, see you then."
}
]
}
Here we are embedding the messages themselves, including their text. Each new message enlarges this document. A conversation can stay active for months or years, so its history is not small, and we cannot predict how large it will get. If the document keeps growing, it can eventually reach the database’s limit on document size.
We also read a conversation differently from an order. Opening an order shows the items in that purchase. Opening a conversation usually shows a recent page of messages, with older messages loaded as the user scrolls.
Recording the relationship on each message
For this chat app, we keep the conversation’s participants and other conversation details in its document, and store messages in a separate messages collection. Each message records which conversation it belongs to:
{
"_id": "message-102",
"conversationId": "conversation-12",
"senderId": "user-18",
"sentAt": "2026-09-28T14:01:00Z",
"text": "Yes, see you then."
}
Adding a message now inserts a document into messages. So as the conversation’s history grows, we get more documents in messages, and no single document gets larger.
How do we find the messages? We could put their identifiers in an array inside the conversation. That array would be much smaller than the full messages, but it would still get one entry for every message. We would also have to keep it consistent with the messages’ conversationId fields.
We do not need that array. Each message already names its conversation. We can retrieve the latest 20 messages by filtering on that identifier and sorting:
db.messages.find({ conversationId: "conversation-12" })
.sort({ sentAt: -1, _id: -1 })
.limit(20)
The -1 specifies descending order, so the newest messages come first. The _id breaks ties when messages have the same timestamp. A compound index supports this read:
db.messages.createIndex({ conversationId: 1, sentAt: -1, _id: -1 })
This is the same direction of relationship we would use with a conversation_id foreign key in a relational messages table: each message identifies its parent, and a query finds the parent’s messages.
So when we decide whether to embed a list, it is not enough to ask whether the records belong together. We also ask how many there may be over time, and whether the application reads the whole group at once or pages through it.