Skip to content

inconsistent behavior of get() methods of PDFOutputDevices  #10

Description

@ashutoshvarma

When calling get() with index out of page range RawImageOutput returns last page's image whereas TextOutput throws a IndexError.

Steps To Reproduce:-

d = x.Document("samples/simple1.pdf")
iout = x.RawImageOutput(d)
tout = x.TextOutput(d)

print(len(d))
print(iout.get(10))            # will return same as iout.get(0)
print(tout.get(10))            # will throw Index Error

Output:-

1
<PIL.Image.Image image mode=RGB size=1275x1651 at 0x7F179F573370>
Traceback (most recent call last):
  File "_test.py", line 15, in <module>
    tout.get(10)
  File "src/pyxpdf/textoutput.pxi", line 268, in pyxpdf.xpdf.TextOutput.get
    cpdef object get(self, int page_no):
  File "src/pyxpdf/textoutput.pxi", line 286, in pyxpdf.xpdf.TextOutput.get
    return self._get_bytes(page_no).decode('UTF-8', errors='ignore')
  File "src/pyxpdf/textoutput.pxi", line 209, in pyxpdf.xpdf.TextOutput._get_bytes
    if self._cache_texts[page_no] == None:
IndexError: list index out of range

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingenhancementNew feature or request

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions