dest = Path('tmp')
url = 'https://s3.amazonaws.com/fast-ai-sample/mnist_tiny.tgz'Core API
Helpers
Most users should use FastDownload rather than calling these helpers.
download_url
def download_url(
url, dest:NoneType=None, timeout:NoneType=None, show_progress:bool=True
):Download url to dest and show progress
download_url shows a progress bar by default. It downloads to a .part file and renames it to the destination after the download completes:
dest.mkdir(exist_ok=True)
fpath = download_url(url, dest)
fpathPath('tmp/mnist_tiny.tgz')
path_stats returns the file size and an MD5 hash of the first megabyte:
path_stats(fpath)(342207, '56143e8f24db90d925d82a5a74141875')
The download_checks.py file containing sizes and hashes will be located next to module:
mod = checks_module(fastdownload)
modPath('git/fastdownload/fastdownload/download_checks.py')
read_checks evaluates the file’s dict literal, and gives an empty dict when there is no file:
assert read_checks({}) == {}check
def check(
fmod, url, fpath
):Check whether size and hash of fpath matches stored data for url or data is missing
check accepts files with no recorded checks for their URL:
update_checks
def update_checks(
fpath, url, fmod
):Store the hash and size of fpath for url in download_checks.py
update_checks records the current size and hash for url, rewriting the whole file:
if mod.exists(): mod.unlink()
update_checks(fpath, url, mod)
read_checks(mod){'https://s3.amazonaws.com/fast-ai-sample/mnist_tiny.tgz': (342207,
'56143e8f24db90d925d82a5a74141875')}
download_and_check
def download_and_check(
url, fpath, fmod, force
):Download url to fpath, unless exists and check fails and not force
Parallel notebook tests can request the same dataset in multiple processes. get holds a lock file beside the archive during download and extraction. Other calls to get wait for the lock, then reuse the extracted data.
FastDownload
def FastDownload(
cfg:NoneType=None, base:str='~/.fastdownload', archive:NoneType=None, data:NoneType=None, module:NoneType=None
):d = FastDownload(module=fastdownload)
d.modulePath('git/fastdownload/fastdownload/download_checks.py')
The config.ini file will be created (if it doesn’t exist) in {base}/config.ini:
d.cfg.config_filePath('.fastdownload/config.ini')
print(d.cfg.config_file.read_text())[DEFAULT]
data = /home/jhoward/.fastdownload/data
archive = /home/jhoward/.fastdownload/archive
FastDownload.download
def download(
url, force:bool=False
):Download url to archive path, unless exists and self.check fails and not force
If there is no stored hash and size for url, or the size and hash matches the stored checks, then download will only download the URL if the destination file does not exist. The destination path will be retured.
if d.module.exists(): d.module.unlink()
arch = d.download(url)
archPath('.fastdownload/archive/mnist_tiny.tgz')
d.update(url)
eval(d.module.read_text()){'https://s3.amazonaws.com/fast-ai-sample/mnist_tiny.tgz': (342207,
'56143e8f24db90d925d82a5a74141875')}
Calling download will now just return the existing file, since the checks match:
d.download(url)Path('.fastdownload/archive/mnist_tiny.tgz')
If the checks file doesn’t match the size or hash of the archive, then a new copy of the file will be downloaded.
FastDownload.extract
def extract(
url, extract_key:str='data', force:bool=False
):Extract archive already downloaded from url, overwriting existing if force
extr = d.extract(url, force=True)
extrPath('.fastdownload/data/mnist_tiny')
extr.ls()(#5) [Path('.fastdownload/data/mnist_tiny/models'),Path('.fastdownload/data/mnist_tiny/train'),Path('.fastdownload/data/mnist_tiny/labels.csv'),Path('.fastdownload/data/mnist_tiny/valid'),Path('.fastdownload/data/mnist_tiny/test')]
Pass extract_key to use a key other than data from your config file when selecting an archive extraction location:
d.cfg['model_path'] = 'models'
d.extract(url, extract_key='model_path')Path('.fastdownload/models/mnist_tiny')
FastDownload.rm
def rm(
url, rm_arch:bool=True, rm_data:bool=True, extract_key:str='data'
):Delete downloaded archive and extracted data for url
d.rm(url)
extr.exists(),arch.exists()(False, False)
get combines the two steps, and is what callers normally use:
FastDownload.get
def get(
url, extract_key:str='data', force:bool=False
):Download and extract url, overwriting existing if force
res = d.get(url)
res,extr.exists()(Path('.fastdownload/data/mnist_tiny'), True)
If the archive doesn’t exist, but the extracted data does, then the archive is not downloaded again.
d.rm(url, rm_data=False)
res = d.get(url)
res,extr.exists()(Path('.fastdownload/data/mnist_tiny'), True)
extract_key works the same way as in FastDownload.extract:
res = d.get(url, extract_key='model_path')
res,res.exists()(Path('.fastdownload/models/mnist_tiny'), True)