Standard-compliant `split` implementation #98

ashvardanian · 2024-02-20T05:55:38Z

The 4c738ea commit introduces a prototype for StringZilla-based Command-Line toolkit, including the split utility replacement. The original prototype suggests a 4x performance improvement opportunity, but it can't currently handle multiple inputs and flags. Those must be easy to add in cli/split.py purely in the Python layer. We are aiming to match the following specification:

$ split --help
Usage: split [OPTION]... [FILE [PREFIX]]
Output pieces of FILE to PREFIXaa, PREFIXab, ...;
default size is 1000 lines, and default PREFIX is 'x'.

With no FILE, or when FILE is -, read standard input.

Mandatory arguments to long options are mandatory for short options too.
  -a, --suffix-length=N   generate suffixes of length N (default 2)
      --additional-suffix=SUFFIX  append an additional SUFFIX to file names
  -b, --bytes=SIZE        put SIZE bytes per output file
  -C, --line-bytes=SIZE   put at most SIZE bytes of records per output file
  -d                      use numeric suffixes starting at 0, not alphabetic
      --numeric-suffixes[=FROM]  same as -d, but allow setting the start value
  -x                      use hex suffixes starting at 0, not alphabetic
      --hex-suffixes[=FROM]  same as -x, but allow setting the start value
  -e, --elide-empty-files  do not generate empty output files with '-n'
      --filter=COMMAND    write to shell COMMAND; file name is $FILE
  -l, --lines=NUMBER      put NUMBER lines/records per output file
  -n, --number=CHUNKS     generate CHUNKS output files; see explanation below
  -t, --separator=SEP     use SEP instead of newline as the record separator;
                            '\0' (zero) specifies the NUL character
  -u, --unbuffered        immediately copy input to output with '-n r/...'
      --verbose           print a diagnostic just before each
                            output file is opened
      --help     display this help and exit
      --version  output version information and exit

The SIZE argument is an integer and optional unit (example: 10K is 10*1024).
Units are K,M,G,T,P,E,Z,Y (powers of 1024) or KB,MB,... (powers of 1000).
Binary prefixes can be used, too: KiB=K, MiB=M, and so on.

CHUNKS may be:
  N       split into N files based on size of input
  K/N     output Kth of N to stdout
  l/N     split into N files without splitting lines/records
  l/K/N   output Kth of N to stdout without splitting lines/records
  r/N     like 'l' but use round robin distribution
  r/K/N   likewise but only output Kth of N to stdout

GNU coreutils online help: <https://www.gnu.org/software/coreutils/>
Report any translation bugs to <https://translationproject.org/team/>
Full documentation <https://www.gnu.org/software/coreutils/split>
or available locally via: info '(coreutils) split invocation'

The text was updated successfully, but these errors were encountered:

Added flags and file input functionalities to split. Updated usage commands in the README. First progress on #97 and #98, but still WIP.

ashvardanian added good first issue Good for newcomers python labels Feb 20, 2024

ashvardanian pushed a commit that referenced this issue Feb 21, 2024

Improve ws and split CLI (#99)

c878caf

Added flags and file input functionalities to split. Updated usage commands in the README. First progress on #97 and #98, but still WIP.

ashvardanian mentioned this issue Feb 21, 2024

Standard-compliant ws and split implementation (Issue 97 98) #99

Merged

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Standard-compliant `split` implementation #98

Standard-compliant `split` implementation #98

ashvardanian commented Feb 20, 2024

Standard-compliant split implementation #98

Standard-compliant split implementation #98

Comments

ashvardanian commented Feb 20, 2024

Standard-compliant `split` implementation #98

Standard-compliant `split` implementation #98