# Weaviate Spark Connector & Async Indexing

**URL:** <https://forum.weaviate.io/t/weaviate-spark-connector-async-indexing/22063>\
**Category:** Support\
**Tags:** integration, developer-experience, technical\
**Created:** [October 19, 2025, 7:18am UTC](https://forum.weaviate.io/t/weaviate-spark-connector-async-indexing/22063 "2025-10-19T07:18:07Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Arindom\_Bora](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.weaviate.io/arindom_bora/32/3737_2.png) [@Arindom\_Bora](https://forum.weaviate.io/u/Arindom_Bora)\
**Post date:** [October 19, 2025, 7:18am UTC](https://forum.weaviate.io/t/weaviate-spark-connector-async-indexing/22063/1 "2025-10-19T07:18:07Z")

</div>

### Description

We have been indexing into weaviate using the spark connector 1.3.2 batch writes. Recently, we came across weaviate’s async indexing capability. When we turned it on using the ENV  
ASYNC\_INDEXING = true & while doing batch writes, we still found the increase in number of vector indexes in the indexed collection to be increasing linearly. We also observed the queue size was increasing as well.

Can someone please explain how weaviate’s async indexing capabilites perform or behave with the spark batch write or is it better to fallback to python client batch write with dynamic server load handling?

---

<div class="post-metadata">

**Author:** ![Shahin](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.weaviate.io/shahin/32/7944_2.png) [@Shahin](https://forum.weaviate.io/u/Shahin)\
**Post date:** [October 20, 2025, 9:09am UTC](https://forum.weaviate.io/t/weaviate-spark-connector-async-indexing/22063/2 "2025-10-20T09:09:13Z")

</div>

Good morning @Arindom_Bora,

Weaviate’s async indexing will for sure put objects in a queue for background indexing however there will be a delay before objects are available for vector search. The queue size reflects objects waiting to be indexed, and it’s normal for this to grow during high imports.

Regarding the Spark connector, I have not come across a proof either bad or good with ASYNC\_INDEXING. **However, async indexing is a server-side feature and should work regardless of the client** , as long as objects are being written to Weaviate.

The main difference with the Python client is that it supports batching features, server-side automatic batching, which dynamically adjusts batch sizes based on server feedback for optimal throughput and stability

> **[Weaviate 1.33 Release | Weaviate](https://weaviate.io/blog/weaviate-1-33-release#server-side-batch-imports)**
>
> 1.33 brings compression by default for optimal resource utilization, powerful 1-bit rotational quantization (RQ), streamlined server-side batch imports, enhanced OIDC group management, and collection aliases become generally available (GA).

> **[Data import | Weaviate Documentation](https://docs.weaviate.io/weaviate/concepts/data-import#server-side-batching)**
>
> Theoretical explanation of client-side and server-side batch imports.

If you need more control or feedback-driven batching, consider the Python client.

Best regards,

**Mohamed Shahin**  
Weaviate Support Engineer  
(Ireland, UTC±00:00/+01:00)

---

<div class="post-metadata">

**Author:** ![Arindom\_Bora](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.weaviate.io/arindom_bora/32/3737_2.png) [@Arindom\_Bora](https://forum.weaviate.io/u/Arindom_Bora)\
**Post date:** [October 20, 2025, 11:41am UTC](https://forum.weaviate.io/t/weaviate-spark-connector-async-indexing/22063/3 "2025-10-20T11:41:24Z")

</div>

Thank You @Shahin . Let me try the python client as well
