# Return "unique file" when search large documents

**URL:** <https://forum.weaviate.io/t/return-unique-file-when-search-large-documents/163>\
**Category:** General\
**Created:** [June 10, 2023, 8:49am UTC](https://forum.weaviate.io/t/return-unique-file-when-search-large-documents/163 "2023-06-10T08:49:00Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Viet\_Tran](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.weaviate.io/viet_tran/32/143_2.png) [@Viet\_Tran](https://forum.weaviate.io/u/Viet_Tran)\
**Post date:** [June 10, 2023, 8:49am UTC](https://forum.weaviate.io/t/return-unique-file-when-search-large-documents/163/1 "2023-06-10T08:49:00Z")

</div>

I’m exploring Weaviate to provide semantic search to find content of attachments in an enterprise application.

I currently chop up each attachment file (usually have a few dozen or a few hundred pages) into chunks. The chunk record would have other meta info like page number, file name, owner application and record id.

Currently when querying using “with\_near\_text”, it usually returns multiple different chunks from the same file.

I’d like it to return one chunk per document (file name) or one chunk per unique application / record Id pair.

I’m wondering is it possible?

---

<div class="post-metadata">

**Author:** ![Viet\_Tran](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.weaviate.io/viet_tran/32/143_2.png) [@Viet\_Tran](https://forum.weaviate.io/u/Viet_Tran)\
**Post date:** [June 12, 2023, 2:55am UTC](https://forum.weaviate.io/t/return-unique-file-when-search-large-documents/163/2 "2023-06-12T02:55:35Z")

</div>

After going through the documentation, my solution is to have two Classes:

- DocumentChunk - properties: text, pageNumber, document (reference to Document class)
- Document - properties: name, summary, ownerId, ownerTable, appName

Query:

```graphql
{
  Get {
    DocumentChunk (
      nearText: {
        concepts: ["overpressure"]
      },
      groupBy: {
        path: ["document"]
        groups: 3,
        objectsPerGroup: 2
      },
    ) {
      text
      pageNumber
      document {
        ...on Document {
          name
        }
      }
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![jphwang](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.weaviate.io/jphwang/32/38_2.png) [@jphwang](https://forum.weaviate.io/u/jphwang)\
**Post date:** [June 12, 2023, 8:51pm UTC](https://forum.weaviate.io/t/return-unique-file-when-search-large-documents/163/3 "2023-06-12T20:51:28Z")

</div>

Hi @Viet_Tran - yeah I think grouping by the parent document makes sense.

Did you know you can group by [the cross-referenced property](https://weaviate.io/developers/weaviate/search/similarity#group-results-by-a-property-or-cross-reference)?

So, depending on what your x-ref is called, you can use replace “document” here with the cross-referenced property.
