# Vectorize big amounts of Data locally

**URL:** <https://forum.weaviate.io/t/vectorize-big-amounts-of-data-locally/2625>\
**Category:** Support\
**Created:** [June 5, 2024, 2:24pm UTC](https://forum.weaviate.io/t/vectorize-big-amounts-of-data-locally/2625 "2024-06-05T14:24:18Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![daza-science](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.weaviate.io/daza-science/32/1679_2.png) [@daza-science](https://forum.weaviate.io/u/daza-science)\
**Post date:** [June 5, 2024, 2:24pm UTC](https://forum.weaviate.io/t/vectorize-big-amounts-of-data-locally/2625/1 "2024-06-05T14:24:18Z")

</div>

I am trying to create a local weaviate database where I store document chunks (which I call passages) and their related vectorized form (using any huggingFace model). I plan to store millions of passages.

I understand I cannot use the HF\_API because there are way too many passages and I do not have a paid account. I also understand how to insert manually the vector, using a local downloaded version from HF, with something like this

```python
with collection_obj.batch.dynamic() as batch:
    for p in corpus.passages:
        props = p.metadata
        props['passage'] = p.text
        batch.add_object(
        properties=props,
        uuid=generate_uuid5(props),
        vector=self.model.embed_passage(p)
         )

```

This works well, but it is veeeery slow! I don’t quite get the weaviate batchin system because I feel that what is happening is that I am embedding one passage at a time… isn’t there a way to parallelize also the embedding step instead of iteratively embedding one passage at a time inside the batch? I cannot pre-vectorize the whole corpus beforehand because, again, there are too many vectors, and I run out of memory.

What is the advice for SCALING UP the “manual vectorization” with weaviate? Beyond the usual tutorial examples…

Thanks for the advice!

---

<div class="post-metadata">

**Author:** ![DudaNogueira](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.weaviate.io/dudanogueira/32/7846_2.png) [@DudaNogueira](https://forum.weaviate.io/u/DudaNogueira)\
**Post date:** [June 7, 2024, 6:18pm UTC](https://forum.weaviate.io/t/vectorize-big-amounts-of-data-locally/2625/2 "2024-06-07T18:18:41Z")

</div>

hi @daza-science !! Welcome to our community! 🤗

Yes! Vectorizing a lot of data concurrently will require beeffy hardware.

Apart from enabling GPU/CUDA, it is about balancing the load with mode nodes, AFAIK.

There are ways of running your own models on _hardware as a service_. We recently published one article on exactly that:

> **[Embed and Search Text at Scale with Modal and Weaviate | Weaviate - Vector...](https://weaviate.io/blog/modal-and-weaviate)**
>
> Learn how to vectorize ~50 million objects and ingest into Weaviate using Modal!

One project/community that is interesting to look at, also, is Ollama:

> **[Ollama](https://ollama.com/)**
>
> Get up and running with large language models.

That on running your own models 🙂

We have a new recipe on using Weaviate with ollama:

> **[recipes/weaviate-features/generative-search at main · weaviate/recipes](https://github.com/weaviate/recipes/tree/main/weaviate-features/generative-search)**
>
> This repository shares end-to-end notebooks on how to use various features and integrations with Weaviate at the core! - weaviate/recipes

Let me know if this helps 🙂
