Andrew Kane
3eef1ff5c2
Removed type-specific code from HNSW [skip ci]
2024-04-24 14:53:45 -07:00
Heikki Linnakangas
b8bdf317f0
Add comment to 'unused' fields
...
I just guessed that these exist for future extendability.
2024-04-24 13:05:02 -07:00
Andrew Kane
78e5bcf229
Switched to 0-based numbering for sparsevec on-disk format
2024-04-24 12:51:24 -07:00
Andrew Kane
4d21eea6f1
Updated comments [skip ci]
2024-04-24 11:27:09 -07:00
Andrew Kane
03ca9adc4c
Added comments [skip ci]
2024-04-24 11:26:05 -07:00
Andrew Kane
d244a040e1
Increased max sparsevec dimensions to 1B [skip ci]
2024-04-24 11:17:25 -07:00
Andrew Kane
c3448a25e2
Improved error messages for sparsevec input
2024-04-24 11:12:28 -07:00
Andrew Kane
053ce2ddae
Improved CI for Windows [skip ci]
2024-04-24 10:22:31 -07:00
Andrew Kane
24c1b51099
Added comment [skip ci]
2024-04-24 10:13:50 -07:00
Andrew Kane
9696835a19
Improved tests for sparsevec input [skip ci]
2024-04-24 09:58:27 -07:00
Andrew Kane
b2a5259607
Switched to strtoint for sparsevec input
2024-04-24 09:56:09 -07:00
Andrew Kane
c198fd58ee
Added more tests for subvector function [skip ci]
2024-04-24 01:31:50 -07:00
Andrew Kane
8c408759dc
Added more tests for subvector function [skip ci]
2024-04-24 01:28:25 -07:00
Heikki Linnakangas
14b351bc92
Fix integer overflow in subvector() function ( #530 )
...
`end = start + count` can overflow if `start` is very large. That
leads to a segfault later in the function. Add test case for it.
2024-04-24 01:20:16 -07:00
Andrew Kane
ad3f811fa3
Use VARSIZE_ANY instead of itemsize to avoid uninitialized bytes
2024-04-23 23:52:02 -07:00
Andrew Kane
281a74f54e
Improved consistency of sparsevec_l1_distance with vector [skip ci]
2024-04-23 21:24:02 -07:00
Andrew Kane
034713c803
Improved consistency with vector [skip ci]
2024-04-23 21:13:00 -07:00
Andrew Kane
ed2e460f00
Improved consistency with vector [skip ci]
2024-04-23 21:11:27 -07:00
Andrew Kane
d136615874
Improved test [skip ci]
2024-04-23 20:42:30 -07:00
Andrew Kane
d70b160e0a
Improved test [skip ci]
2024-04-23 20:41:11 -07:00
Andrew Kane
d1affcc667
Improved tests for l2_norm [skip ci]
2024-04-23 20:38:22 -07:00
Andrew Kane
158481ff2a
Improved tests for sparsevec distance functions [skip ci]
2024-04-23 20:29:04 -07:00
Andrew Kane
794bbaecc7
Removed padding for sparsevec - #529
...
Co-authored-by: Heikki Linnakangas <heikki.linnakangas@iki.fi >
2024-04-23 20:07:24 -07:00
Andrew Kane
8eddcfbd1d
Increased max sparsevec dimensions to 1M [skip ci]
2024-04-23 17:47:11 -07:00
Andrew Kane
b609c343b4
Moved type-specific code to separate functions
2024-04-23 16:32:10 -07:00
Andrew Kane
bbfb3f200a
DRY code for sorting vector arrays [skip ci]
2024-04-23 15:59:42 -07:00
Andrew Kane
99d367edc0
Improved code [skip ci]
2024-04-23 15:53:12 -07:00
Andrew Kane
991743786a
Set length for newCenters and aggCenters [skip ci]
2024-04-23 15:47:04 -07:00
Andrew Kane
60ceaea4f2
Added safety check to NormCenters [skip ci]
2024-04-23 15:43:04 -07:00
Andrew Kane
9cd789fe06
Switched to support function for normalizing centers for k-means
2024-04-23 15:39:58 -07:00
Andrew Kane
0da6213a60
Moved type lookup to support functions - #527
2024-04-23 13:02:47 -07:00
Heikki Linnakangas
d1b83991af
Forbid zero values in sparsevec's binary input function ( #528 )
...
The text input function simply left out any zero values, but the
binary input function did not. That's problematic because you end up
with an "unnormalized" sparse vector, which behaves in weird ways. At
least sparsevec_cmp_internal() expects both inputs to not contain
zeros.
The binary send function never produces such zero values, but an
external tool could. Or to test, you can use COPY TO (FORMAT BINARY),
use a hex editor to edit one of the values to be zero, and copy it
back with COPY FROM (FORMAT BINARY).
2024-04-23 09:13:53 -07:00
Andrew Kane
6c247a38d3
Updated readme [skip ci]
2024-04-22 21:56:05 -07:00
Andrew Kane
6639cde19d
Updated readme [skip ci]
2024-04-22 21:32:59 -07:00
Andrew Kane
bd409f0c6a
Moved HnswGetType call [skip ci]
2024-04-22 19:22:09 -07:00
Andrew Kane
1994fd003a
Removed unneeded headers [skip ci]
2024-04-22 19:10:50 -07:00
Andrew Kane
bd62561a19
Added support function for l2_normalize to ivfflat
2024-04-22 19:06:06 -07:00
Andrew Kane
f14c21748b
Added support function for l2_normalize [skip ci]
2024-04-22 18:36:47 -07:00
Andrew Kane
2b77005610
Removed type-specific code from ivfscan
2024-04-22 18:12:18 -07:00
Andrew Kane
e884b3aa69
Added opclasses to readme [skip ci]
2024-04-22 16:29:34 -07:00
Andrew Kane
ab71c12a28
Added comments on dispatching [skip ci]
2024-04-22 16:18:57 -07:00
Andrew Kane
1804c63e27
Added more tests for vector distance functions [skip ci]
2024-04-22 15:53:13 -07:00
Andrew Kane
4e6aa2f0c1
Added DISABLE_DISPATCH option [skip ci]
2024-04-22 15:43:07 -07:00
Andrew Kane
40e86251c3
Added VECTOR_TARGET_CLONES to VectorL1Distance [skip ci]
2024-04-22 15:15:57 -07:00
Andrew Kane
0c9ae4b187
Added CPU dispatching for L1 distance for halfvec
2024-04-22 15:02:17 -07:00
Andrew Kane
d83af48e70
Improved tests for halfvec l1_distance [skip ci]
2024-04-22 14:43:54 -07:00
Andrew Kane
b2f7dad8a7
Removed support for L1 distance and Jaccard distance from ivfflat due to non-optimal clustering
2024-04-22 14:11:29 -07:00
Andrew Kane
881fbc15ef
Added L1 distance operator to docs [skip ci]
2024-04-22 13:22:28 -07:00
Andrew Kane
f9941c2992
Moved L1 distance to halfutils [skip ci]
2024-04-22 13:19:42 -07:00
Andrew Kane
f9c071a761
Improved tests for L1 distance with halfvec
2024-04-22 13:14:45 -07:00