sij/khoj

mirror of https://github.com/khoj-ai/khoj.git synced 2024-12-19 10:57:45 +00:00

Author	SHA1	Message	Date
Debanjum Singh Solanky	3dd69f7505	Add Upgrade instructions for Obsidian, Emacs to main Readme	2023-01-04 19:50:26 -03:00
Debanjum Singh Solanky	0e39e0ff71	Add details about the Khoj Obsidian plugin to the main Readme - Add Khoj in Obsidian Demo - Update Interfaces Screenshot to include Obsidian Plugin Screenshot - Update .gitignore to ignore obsidian plugin ignorelist Section the .gitignore for better readability - Update the Setup, Usage instructions to include information about the Obsidian plugin	2023-01-04 18:42:53 -03:00
Debanjum Singh Solanky	cd8b918a55	Add manifest.json, versions.json of Obsidian plugin to project root - Obsidian provides limited support for plugins in larger repositories. Currently, it does not have a way to specify the directory of a plugin So it expects the plugins `manifest.json' and `versions.json' to be at project root - While this unnecessarily litters the codebase. It is the (current) required tradeoff for keeping the core plugins in a mono repo	2023-01-04 18:28:16 -03:00
Debanjum Singh Solanky	66ccd0c970	Create Obsidian plugin for Khoj - Features - Search using Khoj from within the Obsidian app Allow Natural language search on your (markdown) notes in Obsidian Vault - Show search results as rendered (instead of raw) Markdown Improve legibility of the results - Jump to selected note from search result in Khoj search modal Simplify seeing result within its original note context - Automatically configure khoj to index markdown files in current vault Reduce khoj setup steps for plugin users by using reasonable defaults - Code updates the markdown config in khoj.yml and triggers index update - It can be configured by user in khoj plugin settings, if required - Add Demo and detailed Readme for the Obsidian plugin Ease setup and usage. Give context about capabilities - Miscellaneous - Trying keep a mono repo until the Khoj project is mature enough to reduce maintainance burden	2023-01-04 18:28:16 -03:00
Debanjum Singh Solanky	e5ef7789fc	Add screenshot of Khoj as PWA on Android Homescreen to Readme	2023-01-04 15:47:08 -03:00
Debanjum Singh Solanky	feddb6ce62	Add start_url to khoj webmanifest to show Khoj as PWA on Chrome	2023-01-04 13:37:56 -03:00
Debanjum Singh Solanky	5ca60a2df7	Add How to Access Khoj on Mobile instructions to Readme	2023-01-04 13:37:40 -03:00
Debanjum Singh Solanky	3dee1aed9e	Create /config/data/default API endpoint to serve default khoj config This can ease configuring khoj from the different interfaces - Don't need to know all the (default) config used by khoj. - Just get default config by calling the above API endpoint. - Then modify desired portions and call POST /api/config/data to configure khoj.	2023-01-03 21:52:34 -03:00
Debanjum Singh Solanky	ce945f7a90	Configure processors too on calling /update API - Previously only search was being reconfigured - But Processors are configured on app start too - Match that behavior on calling /update API	2023-01-03 21:51:02 -03:00
Debanjum Singh Solanky	9d31988f42	Allow starting khoj in non-GUI mode without config file instantiated - Start khoj server (in non-GUI mode) without needing config file already instantiated. - But throw warning to configure khoj to use it - This allows plugins to configure the app via the /config/data APIs - To be used by the Khoj obsidian plugin to configure markdown content in khoj	2023-01-03 21:36:59 -03:00
Debanjum Singh Solanky	52664dd96c	Allow recursive glob pattern (**) to add files to search index - Simplify configuring files to index For Obsidian/Org-Roam type systems with lots of small files in khoj.yml using `input-filter'	2023-01-03 01:32:58 -03:00
Debanjum Singh Solanky	152e5f1661	Return the file of each search result in response - Useful for enabling jump to note functionality in interfaces - It will be used in the Khoj plugin for Obsidian	2023-01-03 01:25:34 -03:00
Debanjum	fe1398401d	Automatically update search index hourly - `c535953` Update index automatically in non GUI mode too - `701d92e` Lock the index before updating it via API or Scheduler - `3b0783a` Automate updating embeddings, search index on a hourly schedule Resolves #106	2023-01-02 00:37:59 +00:00
Debanjum Singh Solanky	c535953915	Update index automatically in non GUI mode too - Poll scheduler every minute using threading.Timer - Use 60 seconds polling interval to avoid fork bombing - Schedule next via the same poll scheduler - Allow clean program interrupt by running scheduler in daemon mode	2023-01-01 21:03:19 -03:00
Debanjum Singh Solanky	701d92e17b	Lock the index before updating it via API or Scheduler - There are 3 paths to updating/setting the index (stored in state.model) - App start - API - Scheduler - Put all updates to the index behind a lock. As multiple updates path that could (potentially) run at the same time (via API or Scheduler)	2023-01-01 17:09:36 -03:00
Debanjum Singh Solanky	3b0783aab9	Automate updating embeddings, search index on a hourly schedule - Use the schedule pypi package - Use QTimer to poll schedule.run_pending() regularly for jobs to run	2023-01-01 17:09:36 -03:00
Debanjum Singh Solanky	a58c243bc0	Document using Word, Date and File Query Filter in Readme	2022-12-26 16:12:49 -03:00
Debanjum	06c25682c9	Split text entries by max tokens supported by ML models ### Background There is a limit to the maximum input tokens (words) that an ML model can encode into an embedding vector. For the models used for text search in khoj, a max token size of 256 words is appropriate [1](https://huggingface.co/sentence-transformers/multi-qa-MiniLM-L6-cos-v1#:~:text=model%20was%20just%20trained%20on%20input%20text%20up%20to%20250%20word%20pieces),[2](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2#:~:text=input%20text%20longer%20than%20256%20word%20pieces%20is%20truncated) ### Issue Until now entries exceeding max token size would silently get truncated during embedding generation. So the truncated portion of the entries would be ignored when matching queries with entries This would degrade the quality of the results ### Fix - `e057c8e` Add method to split entries by specified max tokens limit - Split entries by max tokens while converting [Org](https://github.com/debanjum/khoj/commit/c79919b), [Markdown](https://github.com/debanjum/khoj/commit/f209e30) and [Beancount](https://github.com/debanjum/khoj/commit/17fa123) entries to JSONL - `b283650` Deduplicate results for user query by raw text before returning results ### Results - The quality of the search results should improve - Relevant, long entries should show up in results more often	2022-12-26 18:23:43 +00:00
Debanjum Singh Solanky	17fa123b4e	Split entries by max tokens while converting Beancount entries To JSONL	2022-12-26 15:14:32 -03:00
Debanjum Singh Solanky	f209e30a3b	Split entries by max tokens while converting Markdown entries To JSONL	2022-12-26 13:14:15 -03:00
Debanjum Singh Solanky	24676f95d8	Fix comments, use minimal test case, regenerate test index, merge debug logs - Remove property drawer from test entry for max_words splitting test - Property drawer is not required for the test - Keep minimal test case to reduce chance for confusion	2022-12-25 22:33:04 -03:00
Debanjum Singh Solanky	b283650991	Deduplicate results for user query by raw text before returning results - Required because entries are now split by the max_word count supported by the ML models - This would now result in potentially duplicate hits, entries being returned to user - Do deduplication after ranking to get the top ranked deduplicated results	2022-12-25 21:36:15 -03:00
Debanjum Singh Solanky	53cd2e5605	Regenerate initial model in asymmetric reload test to reduce flakyness - Fix logger message when converting org node to entries - Remove unused import from conftest	2022-12-25 21:36:15 -03:00
Debanjum Singh Solanky	c79919bd68	Split entries by max tokens while converting Org entries To JSONL - Test usage the entry splitting by max tokens in text search	2022-12-25 21:36:00 -03:00
Debanjum Singh Solanky	08dc5e3324	Update instructions in khoj.el to install it from MELPA stable - The instructions suggest installing khoj-assistant via pip install. This installs the latest tagged/release version of khoj - To match that version user should install khoj.el from MELPA stable instead of MELPA	2022-12-23 19:08:38 -03:00
Debanjum Singh Solanky	e057c8e208	Add method to split entries by specified max tokens limit - Issue ML Models truncate entries exceeding some max token limit. This lowers the quality of search results - Fix Split entries by max tokens before indexing. This should improve searching for content in longer entries. - Miscellaneous - Test method to split entries by max tokens	2022-12-23 16:24:04 -03:00
Debanjum Singh Solanky	d3e175370f	Update readme to install khoj.el from MELPA stable unless using pre-release khoj Update readme to ask user to install khoj.el from MELPA when a pre-release version of the main khoj app is installed. Else install khoj.el from MELPA Stable	2022-12-20 23:29:22 -03:00
Debanjum Singh Solanky	cd463c5085	Update Khoj.el Install Instructions on Emacs	2022-12-20 11:06:33 -03:00
Debanjum Singh Solanky	23ca5a2d43	Improve (un-)quoting of funcs used in `khoj--get-enabled-content-types' - Based on melpa package feedback for khoj.el - Verified these changes don't affect behavior of the function	2022-12-19 18:02:23 -03:00
Debanjum Singh Solanky	5db3a67df5	Fix Khoj Emacs package URL in khoj.el	2022-12-14 22:49:19 -03:00
Debanjum Singh Solanky	abad6d5f44	Declare external khoj.el funcs. Remove undefined func warnings on install	2022-12-14 22:36:04 -03:00
Debanjum Singh Solanky	2e5ac5bf22	Bump Khoj to version 0.2.1. It's the current version in development Each push to master creates a development release on pypi	2022-12-04 09:09:34 -03:00
Debanjum Singh Solanky	c52383b11c	Delete stale, unused installation helper script	2022-12-03 13:36:47 -03:00
Debanjum Singh Solanky	676de2372e	Update instructions to install Khoj using Conda - Requires instally PyQT6 using pip as conda doesn't have a package for pyqt6 yet	2022-12-03 13:32:50 -03:00
Debanjum Singh Solanky	1990d09032	Bump khoj version in setup.py, khoj.el to 0.2.0	2022-12-02 14:58:54 -03:00
Debanjum Singh Solanky	a9cfd8b800	Extract hash func for incremental text indexing into separate method	2022-10-26 13:56:58 +05:30
Debanjum Singh Solanky	0de2ff9c97	Add __init__.py to routers directory to register it as a package	2022-10-25 20:40:40 +05:30
Debanjum Singh Solanky	55d2fea9be	Move Custom Formatter class for logger to util.helper module from main.py	2022-10-20 00:32:24 +05:30
Debanjum	5ed17ccbd7	Modularize, Improve API. Formalize Intermediate Text Content Format. Add Type Checking - Improve API Endpoints - `ee65a4f` Merge /reload, /regenerate into single /update API endpoint - `9975497` Type the /search API response to better document the response schema - `0521ea1` Put image score breakdown under `additional` field in search response - Formalize Intermediary Format to Index Text Content - `7e9298f` Use new Text `Entry` class to track text entries in Intermediate Format - `02d9440` Use Base `TextToJsonl` class to standardize `<text>_to_jsonl` processors - Modularize API router code - `e42a38e` Split router code into `web_client`, `api`, `api_beta` routers. Version Khoj API - `d292bdc` Remove API versioning. Premature given current state of the codebase - Miscellaneous - `c467df8` Setup `mypy` for static type checking - `2c54813` Remove unused imports, `embeddings` variable from text search tests	2022-10-19 11:23:04 +00:00
Debanjum Singh Solanky	1c40f97114	Merge branch 'master' of github.com:debanjum/khoj into modularize-api-and-increase-typing - Conflicts: - src/interface/emacs/khoj.el Use our update to `config-url', use their `url-request-method'	2022-10-19 16:46:53 +05:30
Debanjum Singh Solanky	e1b5a87920	Rename Frontend Router to Web Client. Fix logger usage in routers - Use logger in api_beta router instead of print statements - Remove unused logger in web client router	2022-10-19 16:36:48 +05:30
Debanjum	4abd51cb04	Merge pull request #99 from telotortium/method Explicitly set `url-request-method' to GET in khoj.el	2022-10-19 10:31:37 +00:00
Debanjum	74f32eedb8	Remove Exiftool Dependency. Ignore legacy model warning on app start - `bf1ae038cb` Get XMP metadata from image using `Pillow`. Remove `ExifTool` dependency - Pillow library is already used in Khoj and it can extract XMP Metadata from Images - Reduce unmaintained dependencies by using Pillow instead of Exiftool - Pillow is much better maintained than my fork of the Exiftool python package - `c16ae9e344` Ignore "Legacy way to download model" warning for upstream dependency	2022-10-08 14:57:58 +00:00
Debanjum Singh Solanky	c467df8fa3	Setup `mypy' for static type checking	2022-10-08 17:33:13 +03:00
Debanjum Singh Solanky	d292bdcc11	Do not version API. Premature given current state of the codebase - Reason - All clients that currently consume the API are part of Khoj - Any breaking API changes will be fixed in clients immediately - So decoupling client from API is not required - This removes the burden of maintaining muliple versions of the API	2022-10-08 16:32:46 +03:00
Debanjum Singh Solanky	2c548133f3	Remove unused imports, `embeddings' variable from text search tests	2022-10-08 12:06:05 +03:00
Debanjum Singh Solanky	7e9298f315	Use new Text Entry class to track text entries in Intermediate Format - Context - The app maintains all text content in a standard, intermediate format - The intermediate format was loaded, passed around as a dictionary for easier, faster updates to the intermediate format schema initially - The intermediate format is reasonably stable now, given it's usage by all 3 text content types currently implemented - Changes - Concretize text entries into `Entries' class instead of using dictionaries - Code is updated to load, pass around entries as `Entries' objects instead of as dictionaries - `text_search' and `text_to_jsonl' methods are annotated with type hints for the new `Entries' type - Code and Tests referencing entries are updated to use class style access patterns instead of the previous dictionary access patterns - Move `mark_entries_for_update' method into `TextToJsonl' base class - This is a more natural location for the method as it is only (to be) used by `text_to_jsonl' classes - Avoid circular reference issues on importing `Entries' class	2022-10-08 12:06:05 +03:00
Debanjum Singh Solanky	99754970ab	Type the /search API response to better document the response schema - Both Text, Image Search were already giving list of entry, score - This change just concretizes this change and exposes this in the API documentation (i.e OpenAPI, Swagger, Redocs)	2022-10-08 12:06:05 +03:00
Debanjum Singh Solanky	0521ea10d6	Put image score breakdown under `additional' field in search response - Update web, emacs interfaces to consume the scores from new schema	2022-10-08 12:06:01 +03:00
Debanjum Singh Solanky	e42a38e825	Version Khoj API, Update frontends, tests and docs to reflect it - Split router.py into v1.0, beta and frontend (no-prefix) api modules under new router package. Version tag in main.py via prefix - Update frontends to use the versioned api endpoints - Update tests to work with versioned api endpoints - Update docs to mentioned, reference only versioned api endpoints	2022-09-28 20:08:38 +03:00

... 65 66 67 68 69 ...

4082 commits