# \[ERROR\] Batch ConnectionError Exception occurred!

**URL:** <https://forum.weaviate.io/t/error-batch-connectionerror-exception-occurred/616>\
**Category:** Support\
**Created:** [September 5, 2023, 12:17pm UTC](https://forum.weaviate.io/t/error-batch-connectionerror-exception-occurred/616 "2023-09-05T12:17:42Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![kranthi\_kumar](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.weaviate.io/kranthi_kumar/32/433_2.png) [@kranthi\_kumar](https://forum.weaviate.io/u/kranthi_kumar)\
**Post date:** [September 5, 2023, 12:17pm UTC](https://forum.weaviate.io/t/error-batch-connectionerror-exception-occurred/616/1 "2023-09-05T12:17:42Z")

</div>

Hi Team,

Getting below error, on my initial testing and loading the data. Please suggest it any env variables need to change from DB end or in python code.

packages\weaviate\batch\crud\_batch.py:1086: RuntimeWarning: The BatchExecutor was shutdown, most probably when it exited the `with` statement. It will be initialized again. If you are not `batch` in the `with client.batch as batch` please make sure to shut it down when done importing data: `client.batch.shutdown()`. You can start it again using the `client.batch.start()` method.  
warnings.warn(  
2%|█▊ | 430/17766 [05:26\<5:50:01, 1.21s/it][ERROR] Batch ConnectionError Exception occurred! Retrying in 2s. [1/3]  
3%|██▍ | 573/17766 [06:43\<2:41:42, 1.77it/s][ERROR] Batch ConnectionError Exception occurred! Retrying in 2s. [1/3]  
12%|████████▋ | 2069/17766 [22:46\<3:01:53, 1.44it/s][ERROR] Batch ConnectionError Exception occurred! Retrying in 2s. [1/3]  
13%|█████████▌ | 2254/17766 [25:20\<3:47:45, 1.14it/s]Exception in thread batchSizeRefresh:  
Traceback (most recent call last):  
File “C:\Users\x0135069\AppData\Local\Programs\Python\Python310\lib\site-packages\urllib3\connectionpool.py”, line 790, in urlopen  
response = self.\_make\_request(  
File “C:\Users\x0135069\AppData\Local\Programs\Python\Python310\lib\site-packages\urllib3\connectionpool.py”, line 536, in \_make\_request  
response = conn.getresponse()  
File “C:\Users\x0135069\AppData\Local\Programs\Python\Python310\lib\site-packages\urllib3\connection.py”, line 454, in getresponse  
httplib\_response = super().getresponse()  
File “C:\Users\x0135069\AppData\Local\Programs\Python\Python310\lib\http\client.py”, line 1374, in getresponse  
response.begin()  
File “C:\Users\x0135069\AppData\Local\Programs\Python\Python310\lib\http\client.py”, line 318, in begin  
version, status, reason = self.\_read\_status()  
File “C:\Users\x0135069\AppData\Local\Programs\Python\Python310\lib\http\client.py”, line 287, in \_read\_status  
raise RemoteDisconnected(“Remote end closed connection without”  
http.client.RemoteDisconnected: Remote end closed connection without response

During handling of the above exception, another exception occurred:

---

<div class="post-metadata">

**Author:** ![DudaNogueira](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.weaviate.io/dudanogueira/32/7846_2.png) [@DudaNogueira](https://forum.weaviate.io/u/DudaNogueira)\
**Post date:** [September 5, 2023, 3:09pm UTC](https://forum.weaviate.io/t/error-batch-connectionerror-exception-occurred/616/2 "2023-09-05T15:09:42Z")

</div>

Hi!

This looks like some network error.

Can you share how you are importing your data?

---

<div class="post-metadata">

**Author:** ![kranthi\_kumar](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.weaviate.io/kranthi_kumar/32/433_2.png) [@kranthi\_kumar](https://forum.weaviate.io/u/kranthi_kumar)\
**Post date:** [September 6, 2023, 1:03pm UTC](https://forum.weaviate.io/t/error-batch-connectionerror-exception-occurred/616/3 "2023-09-06T13:03:23Z")

</div>

Hi Team,

Please find below conf files we are using.

{  
“cells”: [  
{  
“cell\_type”: “code”,  
“execution\_count”: 16,  
“metadata”: {  
“tags”:   
},  
“outputs”: ,  
“source”: [  
“import os\n”,  
“import openai\n”,  
“import json\n”,  
“from langchain.retrievers import AzureCognitiveSearchRetriever\n”,  
“from langchain.embeddings import OpenAIEmbeddings\n”,  
“from langchain.llms import AzureOpenAI\n”,  
“from langchain.chains import RetrievalQA\n”,  
“from langchain.vectorstores import DocArrayInMemorySearch\n”,  
“from langchain.chat\_models import ChatOpenAI\n”,  
“from langchain.text\_splitter import (RecursiveCharacterTextSplitter,Language)\n”,  
“from langchain.document\_loaders import TextLoader\n”,  
“from langchain.embeddings.openai import OpenAIEmbeddings\n”,  
“from langchain.schema import BaseRetriever\n”,  
“from langchain.vectorstores.azuresearch import AzureSearch\n”,  
“from langchain.vectorstores import Weaviate\n”,  
“import weaviate\n”,  
“from concurrent.futures import ThreadPoolExecutor\n”,  
“from tqdm import tqdm\n”,  
“from openai.error import RateLimitError\n”,  
“import time\n”,  
“import glob\n”,  
“# from docx2python import docx2python\n”,  
“#import PyPDF2\n”,  
“\n”,  
“# from app.openailogging import OpenAILogger\n”,  
“# logger = OpenAILogger(‘promptresponselogs’, ‘prompt-response2.json’)\n”,  
“# os.environ["AZURE\_COGNITIVE\_SEARCH\_SERVICE\_NAME"] = "az-cogsearch-openai-uswst"\n”,  
“# os.environ["AZURE\_COGNITIVE\_SEARCH\_INDEX\_NAME"] = "codebase-pdf-test-uswst"\n”,  
“os.environ["AZURE\_COGNITIVE\_SEARCH\_INDEX\_NAME"] = "EmbeddinganalysisvectorbatchTest"\n”,  
“# os.environ["AZURE\_COGNITIVE\_SEARCH\_API\_KEY"] = ‘uK14Zy2rBdNRPGhhKx36XHxQWpd2CPdDNSjuWQEPdHAzSeDJYFQv’\n”,  
“os.environ[‘OPENAI\_API\_KEY’] = ‘85a8c367d65e435599a4831c2c808013’\n”,  
“os.environ[‘OPENAI\_API\_TYPE’] = "azure"\n”,  
“\n”,  
“os.environ[‘OPENAI\_API\_BASE’] = "[https://openai-useast-dev.openai.azure.com/\“\n”](https://openai-useast-dev.openai.azure.com/%5C%22%5Cn%22),  
“openai.api\_base = os.environ[‘OPENAI\_API\_BASE’]\n”,  
“openai.api\_type = "azure"\n”,  
“openai.api\_key = os.environ[‘OPENAI\_API\_KEY’]\n”,  
“openai.api\_version = "2023-05-15"\n”,  
“model: str = "text-embedding-ada-002"”  
]  
},  
{  
“cell\_type”: “markdown”,  
“metadata”: {},  
“source”: [  
“PLEASE ADD DIFFERENT INDEX NAME TO AVOID OVERWRITING IN THE BELOW CELL”  
]  
},  
{  
“cell\_type”: “code”,  
“execution\_count”: 17,  
“metadata”: {  
“tags”:   
},  
“outputs”: ,  
“source”: [  
“vector\_store\_address: str = "[https://az-cogsearch-openai-uswst.search.windows.net](https://az-cogsearch-openai-uswst.search.windows.net)"\n”,  
“# vector\_store\_password: str = os.environ["AZURE\_COGNITIVE\_SEARCH\_API\_KEY"]\n”,  
“index\_name: str = os.environ["AZURE\_COGNITIVE\_SEARCH\_INDEX\_NAME"]”  
]  
},  
{  
“cell\_type”: “code”,  
“execution\_count”: 18,  
“metadata”: {  
“tags”:   
},  
“outputs”: ,  
“source”: [  
“embeddings: OpenAIEmbeddings = OpenAIEmbeddings(model=model, chunk\_size=1, max\_retries=100)\n”,  
“client = weaviate.Client("http://\<hostname\_replaced\>:8080",timeout\_config=(10,15))\n”,  
“client.batch.configure(batch\_size=50) # Configure batch\n”,  
“vector\_store: Weaviate = Weaviate(\n”,  
" client=client,\n”,  
" index\_name=index\_name,\n",  
" text\_key="content",\n",  
" embedding=embeddings\n",  
“)”  
]  
},  
{  
“cell\_type”: “code”,  
“execution\_count”: 19,  
“metadata”: {  
“tags”:   
},  
“outputs”: ,  
“source”: [  
“# embeddings = OpenAIEmbeddings(deployment="text-embedding-ada-002",openai\_api\_version="2023-03-15-preview", chunk\_size=1)\n”,  
“args ={\n”,  
" "engine": "gpt35turbo",\n",  
“}\n”,  
“llm = ChatOpenAI(temperature=0.3, model\_kwargs=args, verbose=False)\n”,  
“# char\_text\_splitter = RecursiveCharacterTextSplitter.from\_language(language=Language.PYTHON, chunk\_size=2000, chunk\_overlap=5)\n”,  
“char\_text\_splitter = RecursiveCharacterTextSplitter(chunk\_size=2000, chunk\_overlap=5)”  
]  
},  
{  
“cell\_type”: “code”,  
“execution\_count”: 20,  
“metadata”: {  
“tags”:   
},  
“outputs”: ,  
“source”: [  
“def read\_filenames(path,extensions):\n”,  
" files = \n",  
" for extension in extensions:\n",  
" files.extend(glob.glob(path + ‘/\*\*/\*’ + extension, recursive=True))\n",  
" return files"  
]  
},  
{  
“cell\_type”: “code”,  
“execution\_count”: 21,  
“metadata”: {},  
“outputs”: ,  
“source”: [  
“# def read\_files(files):\n”,  
“# file\_contents = {}\n”,  
“# for file in files:\n”,  
“# try:\n”,  
“# pdfReader = PyPDF2.PdfReader(file)\n”,  
“# number\_of\_pages = len(pdfReader.pages)\n”,  
“# file\_content = b’'\n”,  
“# for page\_num in range(0,number\_of\_pages):\n”,  
“# pagehandle = pdfReader.pages[page\_num]\n”,  
“# file\_content += pagehandle.extract\_text().encode(‘UTF-8’)\n”,  
“# filename = "TEXT-" + os.path.basename(file).replace(‘.pdf’, ‘.txt’)\n”,  
“# writing\_path = r’C:\acs\dataset\benchmark\ARK\_DUMP\dump\_all’\n”,  
“# file\_contents[filename] = file\_content\n”,  
“# with open(writing\_path+"\\"+filename, ‘wb’) as f:\n”,  
“# f.write(file\_content)\n”,  
“# except Exception as e:\n”,  
“# print(f"Error processing PDF file{file}:{str(e)}")\n”,  
“# return file\_contents\n”,  
“\n”,  
“# read the files and store them in a list\n”,  
“def read\_files(files):\n”,  
" file\_contents = {}\n",  
" for file in files:\n",  
" with open(file, ‘rb’) as f:\n",  
" writing\_path = r’C:\Software\weaveit\arc\newKbFin\_18k\_snow\newKbFin’\n",  
" file\_contents[os.path.join(writing\_path, os.path.basename(file))] = f.read()\n",  
" return file\_contents\n"  
]  
},  
{  
“cell\_type”: “code”,  
“execution\_count”: 22,  
“metadata”: {},  
“outputs”: [  
{  
“name”: “stdout”,  
“output\_type”: “stream”,  
“text”: [  
" File Extension Count\n",  
“0 .txt 17766\n”  
]  
}  
],  
“source”: [  
“import os \n”,  
“import pandas as pd \n”,  
" \n",  
“# Define the directory you want to start your search from \n”,  
“root\_dir = r’C:\Software\weaveit\arc\newKbFin\_18k\_snow\newKbFin’\n”,  
“\n”,  
" \n",  
“# Define an empty dictionary to store the file extensions and counts \n”,  
“file\_extensions = {} \n”,  
" \n",  
“# Traverse through all the directories and files in the root\_dir \n”,  
“for subdir, dirs, files in os.walk(root\_dir): \n”,  
" for file in files: \n",  
" # Get the file extension \n",  
" file\_ext = os.path.splitext(file)[-1].lower() \n",  
" # Add the file extension to the dictionary and increment the count \n",  
" if file\_ext in file\_extensions: \n",  
" file\_extensions[file\_ext] += 1 \n",  
" else: \n",  
" file\_extensions[file\_ext] = 1 \n",  
" \n",  
“# Convert the dictionary to a pandas dataframe and sort by count in descending order \n”,  
“df = pd.DataFrame(list(file\_extensions.items()), columns=[‘File Extension’, ‘Count’]).sort\_values(‘Count’, ascending=False) \n”,  
" \n",  
“# Print the dataframe \n”,  
“print(df) "  
]  
},  
{  
“cell\_type”: “code”,  
“execution\_count”: 23,  
“metadata”: {},  
“outputs”: [  
{  
“name”: “stdout”,  
“output\_type”: “stream”,  
“text”: [  
“Length of files: 17766\n”  
]  
}  
],  
“source”: [  
“path = r’C:\Software\weaveit\arc\newKbFin\_18k\_snow\newKbFin’\n”,  
“\n”,  
“extensions = [‘.txt’]\n”,  
“files = read\_filenames(path=path, extensions=extensions)\n”,  
“# print(files)\n”,  
“print("Length of files: ", len(files))”  
]  
},  
{  
“cell\_type”: “code”,  
“execution\_count”: 24,  
“metadata”: {},  
“outputs”: ,  
“source”: [  
“contents = read\_files(files)”  
]  
},  
{  
“cell\_type”: “code”,  
“execution\_count”: 25,  
“metadata”: {},  
“outputs”: ,  
“source”: [  
“# contents”  
]  
},  
{  
“cell\_type”: “code”,  
“execution\_count”: null,  
“metadata”: {},  
“outputs”: [  
{  
“name”: “stderr”,  
“output\_type”: “stream”,  
“text”: [  
" 0%| | 15/17766 [00:11\<2:19:30, 2.12it/s]C:\Users\x0135069\AppData\Local\Programs\Python\Python310\lib\site-packages\weaviate\batch\crud\_batch.py:1086: RuntimeWarning: The BatchExecutor was shutdown, most probably when it exited the `with` statement. It will be initialized again. If you are not `batch` in the `with client.batch as batch` please make sure to shut it down when done importing data: `client.batch.shutdown()`. You can start it again using the `client.batch.start()` method.\n”,  
" warnings.warn(\n",  
" 5%|████ | 949/17766 [05:36\<1:30:07, 3.11it/s][ERROR] Batch ConnectionError Exception occurred! Retrying in 2s. [1/3]\n",  
" 14%|██████████▎ | 2449/17766 [14:28\<5:43:21, 1.34s/it]Exception in thread batchSizeRefresh:\n",  
“Traceback (most recent call last):\n”,  
" File "C:\Users\x0135069\AppData\Local\Programs\Python\Python310\lib\site-packages\urllib3\connectionpool.py", line 790, in urlopen\n",  
" response = self.\_make\_request(\n",  
" File "C:\Users\x0135069\AppData\Local\Programs\Python\Python310\lib\site-packages\urllib3\connectionpool.py", line 536, in \_make\_request\n",  
" response = conn.getresponse()\n",  
" File "C:\Users\x0135069\AppData\Local\Programs\Python\Python310\lib\site-packages\urllib3\connection.py", line 454, in getresponse\n",  
" httplib\_response = super().getresponse()\n",  
" File "C:\Users\x0135069\AppData\Local\Programs\Python\Python310\lib\http\client.py", line 1374, in getresponse\n",  
" response.begin()\n",  
" File "C:\Users\x0135069\AppData\Local\Programs\Python\Python310\lib\http\client.py", line 318, in begin\n",  
" version, status, reason = self.\_read\_status()\n",  
" File "C:\Users\x0135069\AppData\Local\Programs\Python\Python310\lib\http\client.py", line 287, in \_read\_status\n",  
" raise RemoteDisconnected("Remote end closed connection without"\n",  
“http.client.RemoteDisconnected: Remote end closed connection without response\n”,  
“\n”,  
“During handling of the above exception, another exception occurred:\n”,

---

<div class="post-metadata">

**Author:** ![DudaNogueira](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.weaviate.io/dudanogueira/32/7846_2.png) [@DudaNogueira](https://forum.weaviate.io/u/DudaNogueira)\
**Post date:** [September 6, 2023, 3:05pm UTC](https://forum.weaviate.io/t/error-batch-connectionerror-exception-occurred/616/4 "2023-09-06T15:05:25Z")

</div>

Is that a jupyter notebook?

---

<div class="post-metadata">

**Author:** ![kranthi\_kumar](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.weaviate.io/kranthi_kumar/32/433_2.png) [@kranthi\_kumar](https://forum.weaviate.io/u/kranthi_kumar)\
**Post date:** [September 6, 2023, 4:03pm UTC](https://forum.weaviate.io/t/error-batch-connectionerror-exception-occurred/616/5 "2023-09-06T16:03:15Z")

</div>

Yes it is Jupiter we are using.

---

<div class="post-metadata">

**Author:** ![kranthi\_kumar](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.weaviate.io/kranthi_kumar/32/433_2.png) [@kranthi\_kumar](https://forum.weaviate.io/u/kranthi_kumar)\
**Post date:** [September 7, 2023, 4:15am UTC](https://forum.weaviate.io/t/error-batch-connectionerror-exception-occurred/616/6 "2023-09-07T04:15:13Z")

</div>

Hi @DudaNogueira ,

Can you please let me know if any issues with Jupiter? or anything need to change in the configurations.

---

<div class="post-metadata">

**Author:** ![DudaNogueira](https://yyz1.discourse-cdn.com/flex027/user_avatar/forum.weaviate.io/dudanogueira/32/7846_2.png) [@DudaNogueira](https://forum.weaviate.io/u/DudaNogueira)\
**Post date:** [September 8, 2023, 2:04pm UTC](https://forum.weaviate.io/t/error-batch-connectionerror-exception-occurred/616/7 "2023-09-08T14:04:10Z")

</div>

Hi Sorry for the delay here 🙂

This notebook is note valid ☹

How big is the data? maybe you’ll need to[throttle it](https://weaviate.io/developers/weaviate/modules/retriever-vectorizer-modules/text2vec-openai#import-throttling).

I can see you are using langchain. For now it is doing at least [one unecessary class check for each added text](https://github.com/langchain-ai/langchain/blob/0672533b3ea10ab49cf17fb25f84fc2c9310700d/libs/langchain/langchain/vectorstores/weaviate.py#L417) this adds some overhead to the import.

What you could try is to ingest the data yourself, following this doc here:

> **[Imports in detail | Weaviate - vector database](https://weaviate.io/developers/weaviate/tutorials/import)**
>
> Overview

Let me know if this helps 🙂
