Skip to content

Releases: Unstructured-IO/unstructured

0.18.1

24 Jun 23:52
3f87946
Compare
Choose a tag to compare

Enhancements

Features

  • Add DocumentData element type This is helpful in scenarios where there is large data that does not make sense to represent across each element in the document.

Fixes

  • The encoding property of the _CsvPartitioningContext is now properly used.

0.17.11-dev1

13 Jun 02:43
5e43e36
Compare
Choose a tag to compare
0.17.11-dev1 Pre-release
Pre-release

What's Changed

New Contributors

Full Changelog: 0.17.2...0.17.11-dev1

0.17.2

20 Mar 16:52
0fa5174
Compare
Choose a tag to compare

Enhancements

  • Add image_url of images in html partitioner <img> tags with non-data content include a new image_url metadata field with the content of the src attribute.

  • Use lxml instead of bs4 to parse hOCR data. lxml is much faster than bs4 given the hOCR data format is regular (garanteed because it is programatically generated)

  • bump numpy to >2. And upgrade paddlepaddle, unstructured-paddleocr, onnx so they are compatible with numpy>2.

Fixes

  • Fix Image in a
    tag is "UncategorizedText" with no .text

What's Changed

Full Changelog: 0.17.0...0.17.2

0.17.0

12 Mar 15:57
2dceac3
Compare
Choose a tag to compare

What's Changed

Full Changelog: 0.16.25...0.17.0

0.16.25

07 Mar 11:17
74b0647
Compare
Choose a tag to compare

0.16.25

Enhancements

Features

Fixes

  • Fixes filetype detection for jsons passed as byte streams - Now it prioritizes magic mimetype prediction over file extension when detecting filetypes

0.16.24

07 Mar 11:17
961c8d5
Compare
Choose a tag to compare

0.16.24

Enhancements

  • Support dynamic partitioner file type registration. Use create_file_type to create new file type that can be handled
    in unstructured and register_partitioner to enable registering your own partitioner for any file type.

  • extract_image_block_types now also works for CamelCase elemenet type names. Previously NarrativeText and similar CamelCase element types can't be extracted using the mentioned parameter in partition. Now figures for those elements can be extracted like Image and Table elements

  • use block matrix to reduce peak memory usage for pdf/image partition.

Features

  • Add JSON elements to HTML converter - Converts JSON elements file into an HTML file.

Fixes

0.16.23

20 Feb 13:31
0df50fe
Compare
Choose a tag to compare

0.16.23

Enhancements

Features

Fixes

  • Fixes detect_filetype when SpooledTemporaryFile is passed. Previously some random name would get assigned to the file and the function raised error.

0.16.22

20 Feb 01:11
147add9
Compare
Choose a tag to compare

0.16.22

Enhancements

Features

Fixes

  • Fix open CVES in and bump dependencies

0.16.21

17 Feb 16:01
3403db1
Compare
Choose a tag to compare

Enhancements

  • Use password to load PDF with all modes

  • use vectorized logic to merge inferred and extracted layouts. Using the new LayoutElements data structure and numpy library to refactor the layout merging logic to improve compute performance as well as making logic more clear

  • Add PDF Miner configuration Now PDF Miner can be configured via pdfminer_line_overlap, pdfminer_word_margin, pdfminer_line_margin and pdfminer_char_margin parameters added to partition method.

Features

Fixes

  • Fix file type detection for NDJSON files NDJSON files were being detected as JSON due to having the same mime-type.

0.16.20

06 Feb 06:12
b10379c
Compare
Choose a tag to compare

0.16.20

Enhancements

Features

Fixes

  • Fix a security issue where rst and org files could read files in the local filesystem. Certain filetypes could 'include' or 'import' local files into their content, allowing partitioning of arbitrary files from the local filesystem. Partitioning of these files is now sandboxed.