sij/khoj

mirror of https://github.com/khoj-ai/khoj.git synced 2024-11-27 17:35:07 +01:00

Author	SHA1	Message	Date
Debanjum Singh Solanky	cbae8b68fb	Add DB migration from making bi_encode configs optional in #834	2024-07-11 16:33:31 +05:30
Debanjum Singh Solanky	3a75838196	Add Keyboard shortcuts to navigate in Khoj Desktop	2024-07-11 16:29:53 +05:30
Debanjum Singh Solanky	6c1861b319	Improve the prompt to generate images with DALLE3 and SD3 - Major - Ask for prompt in prose - Remove seed from SD3 image generation to improve diversity of output for a given prompt Otherwise for conversations with similar sounding prompts, the images would be almost exactly the same. This maybe another indicator of SD3's inability to capture detailed instructions - Consistently use "prompt" wording instead of "query" in improved image generation prompts. Previously a mix of those terms were being used, which could confuse the chat model - Minor - Add day of week to prompt - Remove 2-5 sentence limit on instructions to SD3. It seems to be able to follow longer instructions just with less fidelity than DALLE. And the 2-5 sentence instruction limit wasn't being adhered to - Improve ability to edit, improve the image based on follow-up instructions by the user - Align prompts for DALLE and SD3. Only difference is to wrap text to be rendered in quotes for SD3. This improves it's ability to render requested text. DALLE cannot render text as well or consistently	2024-07-11 16:29:53 +05:30
sabaimran	260aa61818	Remove tests for python3.9	2024-07-09 12:28:11 +05:30
sabaimran	4471c1e37f	Apply mitigations for piling up open connections - Because we're using a FastAPI api framework with a Django ORM, we're running into some interesting conditions around connection pooling and clean-up. We're ending up with a large pile-up of open, stale connections to the DB recurringly when the server has been running for a while. To mitigate this problem, given starlette and django run in different python threads, add a middleware that will go and call the connection clean up method in each of the threads.	2024-07-09 12:22:58 +05:30
Debanjum	0b1b262512	Add system dependencies required by RapidOCR to fix Khoj Docker image (#842 ) - Issue The Khoj docker build would fail with `ImportError: libGL.so.1: cannot open shared object file: No such file or directory`. This was required by the Khoj RapidOCR python package dependency. - Fix A minimal set of system packages have been added to resolve this issue.	2024-07-08 22:16:16 +05:30
kxnarak	43413cd21f	add dependencies required by the RapidOCR python package	2024-07-08 18:26:19 +05:30
sabaimran	037e157648	Fix a variety of links	2024-07-08 16:49:13 +05:30
sabaimran	6b80bb3f37	Add a demo for the khoj mini application, minor updates to other pages, remove out of date demos page	2024-07-08 16:33:47 +05:30
Debanjum Singh Solanky	9e31ebff93	Release Khoj version 1.16.0	2024-07-07 18:26:10 +05:30
Debanjum Singh Solanky	54132efd67	Fix Khoj Obsidian plugin build	2024-07-07 18:26:10 +05:30
Debanjum Singh Solanky	510d9b3a29	Add short keys to open chat menu, new chat, search from Obsidian pane	2024-07-07 17:57:17 +05:30
Debanjum Singh Solanky	3e0c882e27	Transcribe only when keyboard shortcut or button pressed in Obsidian - Transcribe on holding Ctrl+s keyboard shortcut - Transcribe on holding the transcribe button pressed via mouse too - Make the transcribe button robust to inadvertent touches by using timeout - Do not transcribe, trigger auto-send on silences. Silence detection is super rudimentary, just blocks standard emanations by whisper when no speech	2024-07-07 17:57:17 +05:30
sabaimran	0eb000c3ea	Add health checks for the django ORM	2024-07-07 16:11:28 +05:30
Debanjum Singh Solanky	a31cd0dec1	Fix async batch delete of indexed entries	2024-07-06 22:45:26 +05:30
Debanjum	08b379c2ab	Fix, Improve Indexing, Deleting Files (#840 ) ### Fix - Fix degrade in speed when indexing large files - Resolve org-mode indexing bug by splitting current section only once by heading - Improve summarization by fixing formatting of text in indexed files ### Improve - Improve scaling user, admin flows to delete all entries for a user	2024-07-06 19:52:42 +05:30
Debanjum Singh Solanky	4a471979eb	Upgrade sentence-transformer package to version 3.0.1 Add einops dependency for some sentence transformer models like the nomic-embed	2024-07-06 19:35:59 +05:30
Debanjum Singh Solanky	d693baccbc	Make it optional to set the encoder, cross-encoder configs via admin UI	2024-07-06 19:35:59 +05:30
Debanjum Singh Solanky	1baebb8d0e	Identify markdown headings by any whitespace character after ^#+ Previously only markdown headings with space characters after # would be considered a heading. So ^##\t wouldn't be considered a valid heading	2024-07-06 19:35:59 +05:30
Debanjum Singh Solanky	010486fb36	Split current section once by heading to resolve org-mode indexing bug - Split once by heading (=first_non_empty) to extract current section body Otherwise child headings with same prefix as current heading will cause the section split to go into infinite loop - Also add check to prevent getting into recursive loop while trying to split entry into sub sections	2024-07-06 19:35:59 +05:30
Debanjum Singh Solanky	6a135b1ed7	Fix degrade in speed of indexing large files. Improve summarization Adding files to the DB for summarization was slow, buggy in two ways: - We were updating same text of modified files in DB = no of chunks per file times - The `" ".join(file_content)' code was breaking each character in the file content by a space. This formats the original file content incorrectly before storing in the DB Because this code ran in the main file indexing path, it was slowing down file indexing. Knowledge bases with larger files were impacted more strongly	2024-07-06 19:35:59 +05:30
Debanjum Singh Solanky	e6ffb6b52c	Improve scaling user flow to delete all entries - Delete entries by batch to improve efficiency of query at scale - Share code to delete all user entries between it's async, sync methods - Add indicator to show when files being deleted on web config page	2024-07-06 19:35:59 +05:30
Debanjum Singh Solanky	1ab59865b5	Improve scaling admin flow to delete all entries for user	2024-07-06 19:35:59 +05:30
Debanjum	05138cbd0a	Use DOM Scripting, Add CSP to Web config pages. Disable CSP in Obsidian plugin (#834 ) - Add CSP to web config pages. Load phone no. validation js, css from S3 - Construct config page elements on Web via DOM scripting - Disable CSP in Khoj Obsidian as it interferes with Obsidian functionality - Other miscellaneous voice message level improvements (rate limit, listening animation)	2024-07-06 19:30:09 +05:30
Debanjum Singh Solanky	9bdb48807b	Ratelimit text to speech model. Validate share chat url domain - Do not log auth error message on server when Resend setup as Magic links for sign-in are now supported	2024-07-06 12:53:19 +05:30
Debanjum Singh Solanky	b334db0fca	Add CSP to web config pages. Load phone no validation js, css from S3	2024-07-06 12:48:28 +05:30
Debanjum Singh Solanky	2f034f807a	Construct config page elements on Web via DOM scripting. Minimize isage of innerHTML to prevent DOM clobbering and unintended escape by user Input	2024-07-06 12:48:28 +05:30
Debanjum Singh Solanky	69c9e8cc08	Disable CSP in Khoj Obsidian as it interferes with Obsidian functionality The Khoj CSP interferes with other Obsidian features and plugins as CSP is applied page wide. For now chat message sanitization via Dompurify should suffice. Enable CSP when can scope it to only the Khoj Obsidian plugin.	2024-07-05 16:10:08 +05:30
Debanjum Singh Solanky	a353d883a0	Make it optional to set the encoder, cross-encoder configs via admin UI Upgrade sentence-transformer, add einops dependency for some sentence transformer models like nomic	2024-07-05 16:09:30 +05:30
Debanjum Singh Solanky	6d59ad7fc9	Add listening circle animation to speak button in Obsidian plugin Use icon active focus as color of animation button	2024-07-05 14:00:53 +05:30
Debanjum Singh Solanky	516af86575	Fix add, remove of the text to speech loader element in Obsidian	2024-07-04 17:38:45 +05:30
Debanjum Singh Solanky	814aca6d69	Skip summarize when not triggered via slash cmd and can't summarize Maybe better to fallback to non-summarize behavior if summarize intent is just inferred but we can't actually summarize because the single file added to conversation isn't satisfied	2024-07-04 13:31:00 +05:30
Debanjum	4446de00d3	Enable Voice, Keyboard Shortcuts in Khoj Obsidian Plugin (#837 ) - Simplify quick jump between Khoj side pane and main editor view using keyboard shortcuts - Enable voice chat in Obsidian to make interactions with Khoj more seamless	2024-07-04 13:28:29 +05:30
sabaimran	5ea8b16f84	Fix missing method error	2024-07-04 12:08:22 +05:30
sabaimran	d61bddf56c	Fix retrieving image model by prefetching the openai config in the async method	2024-07-04 11:58:33 +05:30
sabaimran	a129b017b9	Fix image generation on server -- use default config when not set by user	2024-07-04 09:13:23 +05:30
sabaimran	6fa2dbc042	Do not use the custom configured max prompt size to send message to anthropic	2024-07-02 21:59:06 +05:30
Debanjum Singh Solanky	afcfc60637	Merge DB migrations post merge of SD3 via API support PR	2024-07-02 17:54:58 +05:30
Debanjum	c015eeb5dd	Improve Online Search: Parallelize Search, Use Jina Reader API by default (#832 ) - Overview Khoj wil be able to do online search out of the box, even for self-hosted users - Default to Jina search, reader API when no Serper.dev, Olostep API keys - Run online searches in parallel to process multiple queries faster - Details - Jina provides a [reader API](https://github.com/jina-ai/reader) for online search and web page reading It requires no API key. This provides a good default to enable online search for self-hosted readers requiring no additional setup. - Jina search API also returns webpage contents with the results, so just use those directly when Jina Search API used instead of trying to read webpages separately. The extract relevant content from webpage step using a chat model is still used from the `read_webpage_and_extract_content' func in this case. - Parse search results from Jina search API into same format as Serper.dev for accurate rendering of online references by clients - Run online searches in parallel with AsyncIO to process multiple queries faster	2024-07-02 17:44:51 +05:30
Debanjum	826c3dc9cc	Enable using Stable Diffusion 3 for Image Generation via API (#830 ) - Support Stable Diffusion 3 via API Server Admin needs to setup model similar to DALLE-3 via Django Admin Panel - Use shorter prompt generator to prompt SD3 to create better images - Allow users to set paint model to use from web client config page	2024-07-02 17:28:50 +05:30
Debanjum Singh Solanky	d5ceff2691	Update tests and documentation with Jina reader API usage and info Update offline, openai chat actor, director tests to not require Serper to run the online command tests Update documentation for self-hosted online search to mention no setup is required by default. But improvements can be made by using Serper.dev or Olostep	2024-07-02 17:19:09 +05:30
Debanjum Singh Solanky	553beae848	No need to set OpenAI API key from environment variable explicitly It is unnecessary as the OpenAI client automatically tries to use API key from OPENAI_API_KEY env var when the api_key field is unset	2024-07-02 17:19:09 +05:30
Debanjum Singh Solanky	a038e4911b	Default to Jina search, reader API when no Serper.dev, Olostep API keys Jina AI provides a search and webpage reader API that doesn't require an API key. This provides a good default to enable online search for self-hosted readers requiring no additional setup. Jina search API also returns webpage contents with the results, so just use those directly when Jina Search API used instead of trying to read webpages separately. The extract relvant content from webpage step using a chat model is still used from the `read_webpage_and_extract_content' func in this case. Parse search results from Jina search API into same format as Serper.dev for accurate rendering of online references by clients	2024-07-02 17:19:08 +05:30
Debanjum Singh Solanky	ff44734774	Run online searches in parallel to process multiple queries faster	2024-07-02 17:19:08 +05:30
Raghav Tirumale	8eccd8a5e4	Support Indexing Images via OCR (#823 ) - Added support for uploading .jpeg, .jpg, and .png files to Khoj from Web, Desktop app - Updating indexer to generate raw text and entries using RapidOCR - Details * added support for indexing images via ocr * fixed pyproject.toml * Update src/khoj/processor/content/images/image_to_entries.py Co-authored-by: Debanjum <debanjum@gmail.com> * Update src/khoj/processor/content/images/image_to_entries.py Co-authored-by: Debanjum <debanjum@gmail.com> * removed redudant try except blocks * updated desktop js file to support image formats * added tests for jpg and png * Fix processing for image to entries files * Update unit tests with working image indexer * Change png test from version verificaition to open-cv verification --------- Co-authored-by: Debanjum <debanjum@gmail.com> Co-authored-by: sabaimran <narmiabas@gmail.com>	2024-07-01 06:00:00 -07:00
Debanjum Singh Solanky	cffc14a46a	Trigger voice chat via keyboard shortcut in Khoj side pane Quickly trigger voice chat from Khoj side pane using Keyboard shortcuts	2024-07-01 18:06:09 +05:30
Debanjum Singh Solanky	3723904512	Toggle jump between Khoj side pane & previous editor via cmd, kbd shortcut Improve quick navigation to, from Khoj side pane using Keyboard shortcut or Obsidian command	2024-07-01 18:05:59 +05:30
Debanjum Singh Solanky	fbb95ca342	Put cursor on chat input when focus on chat view in Obsidian This should improve fluidity of keyboard interactions with Khoj on Obsidian. Open Khoj chat view via keybinding or command pallete and ask question using only the keyboard, with no mouse clicks required	2024-07-01 18:05:55 +05:30
Debanjum Singh Solanky	093e276908	Enable Voice chat in Khoj Obsidian plugin - Automatically carry out voice chats with Khoj from within Obsidian When send voice message, Khoj will auto respond with voice as well - Listen to past Khoj messages as speech - Add circular loading spinner to use while message is being converted to speech	2024-07-01 18:02:28 +05:30
sabaimran	c83b8f2768	Allow just one worker to be the background schedule leader (#836 ) * Add a leader election mechanism to circumvent runtime issues for multiple schedulers - Reduce the load on the DB and risk of issues on the service side by limiting the execution environment to one elected leader at a given time. This one is responsible for managing all of the execution of the jobs, though all workers are capable of adding and removing jobs * Set a max duration for the schedule leader position (12 hrs), add some error if automation not added successfully	2024-06-28 13:13:25 +05:30

1 2 3 4 5 ...

2975 commits